minimaxiH3 Get access
VideoVendor content policy appliesEarly access

数字人口播

One photo. One voice. A presenter.

A portrait and an audio file become a presenter: head movement, blinks, gestures and lip-sync are all driven by the sound. For product clips, intros, or characters who need to speak to camera.

一张脸、一段音频,出一个会说话的人。

Price From $0.072 per second on the cheapest model below.

Talking avatar

from $0.072/sec
Drop a face and an audio track hereor click to choose filesFiles stay in your project. Nothing is trained on.
length
frame
Invite code opens this tool in Chat. No code yet? Join the waitlist — we open the most-requested tools first.

How it works

Three moves.

Upload a portrait

Shoulders up, neutral expression.

Upload audio

Clean voice, little background noise.

Run

Pick a calmer or livelier gesture style if offered.

Illustrative sample from Avatar Omni Human 1.5 — not a before/after of this tool.

You give

Inputs.

  • One portrait
  • Audio (speech or singing)
  • Optional gesture style

You get

Outputs.

  • Talking-head clip, audio embedded

Under the hood

Models behind this tool.

Each model lists its own price and content policy. The spicy route is MiniMax H3; other models follow their vendor’s rules.

FAQ

Short answers.

How long can the audio be?

Typically up to 30–60 seconds per run depending on the model; chain runs for longer scripts.

Can the avatar be a synthetic character?

Yes — generate the portrait with a text-to-image model first.

Is this spicy-route?

No. Avatar models apply their vendor policy; keep the content within it.

How do I get access?

Enter an invite code to open the tool in Chat, or join the waitlist from this page. We open seats in batches and prioritise the tools people ask for most.

Related tools

People also open.

Talking avatar, in Chat.

Invite code opens the tool. No code — join the waitlist; the most-requested tools open first.

Get access Browse models