Avatar Omni Human 1.5
OmniHuman 1.5 is ByteDance's digital-human model that turns a single portrait plus an audio track into a lifelike video of that character speaking or singing, with lip-sync, expressions, and gestures generated straight from the audio.
What it does
Avatar Omni Human 1.5, in practice.
Turn a single portrait photo into a lifelike, lip-synced talking video — driven entirely by an audio clip. Avatar Omni Human 1.5 (OmniHuman) by ByteDance is a state-of-the-art digital-human video generation model. Give it one reference image of a person and an audio track, and it generates a natural, expressive video of that person speaking — with accurate lip sync, head motion, and facial expressions that match the audio.
- Audio-driven lip sync — mouth movements precisely follow the speech in your audio.
- Identity preservation — the generated person stays faithful to your reference image.
- Natural motion — lifelike head pose, blinking, and micro-expressions, not a stiff talking head.
- Multilingual — works with audio in many languages, including Chinese, English, Japanese, Korean, Spanish, and Indonesian.
- High resolution — render output at up to 1080p.
Run Avatar Omni Human 1.5
from $0.18/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
image_url | URL of the reference portrait image — a clear, front-facing face works best | |
audio_url | URL of the driving audio (MP3/WAV). Maximum 60 seconds; 15 seconds or less recommended for best quality. The avatar lip-syncs to this audio. | |
prompt | Optional hint for action / scene / expression (supports Chinese, English, Japanese, Korean, Spanish, Indonesian) | |
output_resolution | Output video resolution (720 or 1080) | 720 1080 |
seed | Random seed; -1 for random | default -1 |
FAQ
Short answers.
How much does Avatar Omni Human 1.5 cost?
Pricing starts at $0.18 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $1.44. Usage is billed per request from your balance — no subscription.
Does Avatar Omni Human 1.5 run uncensored here?
No. ByteDance applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists Avatar Omni Human 1.5 for everything else it does well.
What does Avatar Omni Human 1.5 take as input?
It is a audio-to-video model. OmniHuman 1.5 is ByteDance's digital-human model that turns a single portrait plus an audio track into a lifelike video of that character speaking or singing, with lip-sync, expressions, and gestures generated straight from the audio.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related