HappyHorse-1.1 Reference-to-video
Generates videos from one to nine reference images and a text prompt, supporting 480P, 720P, or 1080P output, flexible aspect ratios, and durations from 3 to 15 seconds.
What it does
HappyHorse-1.1 Reference-to-video, in practice.
Alibaba HappyHorse 1.1 Reference-to-Video creates short video clips from a text prompt and one or more reference images, helping preserve subjects, objects, or visual cues from the supplied images.
- Reference-guided generation: Use 1 to 9 images to guide the people, objects, style, or scene elements in the output.
- Prompt-directed motion: Describe how the referenced elements should move, interact, or appear in the final clip.
- Flexible framing: Supports `16:9`, `9:16`, `1:1`, `4:3`, `3:4`, `4:5`, `5:4`, `9:21`, and `21:9`.
- Two output resolutions: Generate at `720P` or `1080P`.
- Creators producing videos that need to keep a character, product, or prop recognizable.
- Teams building campaign visuals from approved reference material.
Run HappyHorse-1.1 Reference-to-video
from $0.11/secParameters
What you can set.
| Parameter | What it does | Options |
|---|---|---|
prompt | Text prompt describing the desired video and how the reference images should be used. Maximum length is 2500 characters. | |
images | Reference image URLs. Provide 1 to 9 images. Supported formats are JPEG, JPG, PNG, and WEBP. Each image can be up to 20 MB, with the shorter side at least 400 px. | |
resolution | Output video resolution. | 480p 720p 1080p |
ratio | Aspect ratio of the generated video. | 16:9 9:16 1:1 4:3 3:4 4:5 5:4 9:21 21:9 |
duration | Video duration in seconds. | default 5 |
seed | Random seed for video generation. Use -1 for a random seed. | default -1 |
Sample prompt
The prompt behind the sample.
A tiny medieval knight stands alone on a weathered stone path, wearing a worn iron helmet and a dark leather cloak. Gentle wind blows across the scene, causing the cloak to sway softly and small dust particles to drift through the air. The knight slowly raises its head as if sensing a distant call. The camera begins with a close-up of the helmet visor, revealing scratches, rust, and detailed textures, then smoothly pulls back and circles around the character. Soft morning fog rolls across the ground while golden sunlight pierces through the mist, creating cinematic volumetric rays. In the dist
resolution: 1080Pratio: 16:9duration: 5seed: -1FAQ
Short answers.
How much does HappyHorse-1.1 Reference-to-video cost?
Pricing starts at $0.11 per second of video at the base resolution; higher resolutions and longer durations cost more. An 8-second clip at the base tier is about $0.84. Usage is billed per request from your balance — no subscription.
Does HappyHorse-1.1 Reference-to-video run uncensored here?
No. Qwen applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists HappyHorse-1.1 Reference-to-video for everything else it does well.
What does HappyHorse-1.1 Reference-to-video take as input?
It is a text-to-video model. Generates videos from one to nine reference images and a text prompt, supporting 480P, 720P, or 1080P output, flexible aspect ratios, and durations from 3 to 15 seconds.
How do I use it?
Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.
Related