Describe the frame
Subject, wardrobe, light, lens.
文生图 / 首帧
Most good clips start as a still. Generate the frame with the model that suits the look — photoreal, editorial, anime — at 1K to 4K, then send it to image-to-video. Each model lists its per-image price and its content policy.
先出一张对的图,再去让它动。每个模型标了每张价格和内容政策。
How it works
Subject, wardrobe, light, lens.
Seedream for realism, Qwen for text, GPT Image for edits.
The still becomes the locked first frame.

You give
You get
Under the hood
Each model lists its own price and content policy. The spicy route is MiniMax H3; other models follow their vendor’s rules.
FAQ
Image models here follow their vendor policy; the spicy route is the MiniMax H3 video models. Generate a still within policy, then animate it on H3.
Match the video aspect ratio at 1080 px or more on the short edge.
From a few cents at the base tier; see each model page.
Enter an invite code to open the tool in Chat, or join the waitlist from this page. We open seats in batches and prioritise the tools people ask for most.
Related tools
Invite code opens the tool. No code — join the waitlist; the most-requested tools open first.
Early access
Text-to-image, first frames is invite-only for now. Enter a code, or leave your email and we will count the request.