minimaxiH3 Get access
ByteDanceAudio-to-videoVendor content policy applies

Avatar Omni Human 1.5

OmniHuman 1.5 is ByteDance's digital-human model that turns a single portrait plus an audio track into a lifelike video of that character speaking or singing, with lip-sync, expressions, and gestures generated straight from the audio.

$0.18per secondStarting price at the base resolution and quality tier.
$1.44Typical 8-second clip at the base tier. Longer or higher-resolution outputs scale with duration and tier.
Pay per useBilled per request from your balance. No subscription, no minimum.
Sample output

What it does

Avatar Omni Human 1.5, in practice.

Turn a single portrait photo into a lifelike, lip-synced talking video — driven entirely by an audio clip. Avatar Omni Human 1.5 (OmniHuman) by ByteDance is a state-of-the-art digital-human video generation model. Give it one reference image of a person and an audio track, and it generates a natural, expressive video of that person speaking — with accurate lip sync, head motion, and facial expressions that match the audio.

  • Audio-driven lip sync — mouth movements precisely follow the speech in your audio.
  • Identity preservation — the generated person stays faithful to your reference image.
  • Natural motion — lifelike head pose, blinking, and micro-expressions, not a stiff talking head.
  • Multilingual — works with audio in many languages, including Chinese, English, Japanese, Korean, Spanish, and Indonesian.
  • High resolution — render output at up to 1080p.

Run Avatar Omni Human 1.5

from $0.18/sec
Drop audio and a face image hereor click to choose a fileFiles stay in your project. Nothing is trained on.
output_resolution
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
image_urlURL of the reference portrait image — a clear, front-facing face works best
audio_urlURL of the driving audio (MP3/WAV). Maximum 60 seconds; 15 seconds or less recommended for best quality. The avatar lip-syncs to this audio.
promptOptional hint for action / scene / expression (supports Chinese, English, Japanese, Korean, Spanish, Indonesian)
output_resolutionOutput video resolution (720 or 1080)720 1080
seedRandom seed; -1 for randomdefault -1

FAQ

Short answers.

How much does Avatar Omni Human 1.5 cost?

Pricing starts at $0.18 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.44. Usage is billed per request from your balance — no subscription.

Does Avatar Omni Human 1.5 run uncensored here?

No. ByteDance applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Avatar Omni Human 1.5 for everything else it does well.

What does Avatar Omni Human 1.5 take as input?

It is a audio-to-video model. OmniHuman 1.5 is ByteDance's digital-human model that turns a single portrait plus an audio track into a lifelike video of that character speaking or singing, with lip-sync, expressions, and gestures generated straight from the audio.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run Avatar Omni Human 1.5.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models