Create consistent character-driven short videos with multiple reference images
happyhorse-1.1-r2v is the reference-image-to-video model for HappyHorse 1.1, suitable for tasks with existing character, clothing, or prop assets where you want to continue creating short videos in new scenes. It combines ordered reference images with textual shot descriptions, using visual assets to constrain subjects and prompts to arrange actions, environments, and camera work, with a focus on short-video creation requiring high character consistency.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation mode
Reference-image-to-video; input prompt and 1–9 image_urls
Asset references
Use names such as character1 and character2 according to image order
Native output
480P, 720P, 1080P; 24 fps; MP4
API resolutions
720P, 1080P; default 1080P
Video duration
3–15 seconds; default 5 seconds
Available aspect ratios
16:9, 9:16, 1:1, 4:3, 3:4
Task delivery
Supports asynchronous queries and completion callbacks, returning video_url
Native specifications include audio capability and 480P output; this platform entry provides reference-image generation calls at 720P and 1080P.
Core capabilities
Extend character identity with reference images
Unlike describing appearance with text alone, reference images provide clear visual grounding for characters, clothing, and props. You can change environments and actions in the prompt while continuing to use the same character assets, making it suitable for creating short shots of one character across different scenes; consistency is a creative goal, not a guarantee of complete frame-by-frame uniformity.
Assign asset roles in sequence
Images are not unordered attachments: use character1 and character2 to point to corresponding assets, and specify which image provides the subject and which provides clothing or visual style. For example, have character1 walk through grass while adopting the leather and gold-trim style of character2, establishing a clear relationship between assets and shot descriptions.
From generation task to video delivery
Submit a task through POST /happyhorse/videos, explicitly select reference_to_video and this model, then set duration, resolution, and aspect ratio. When long connections are inconvenient, use async or callback_url; after saving task_id, query the results or receive completion notifications to obtain the video link for downloading and further editing.
Applicable Scenarios
Character Series Clips
Input reference images of the same character, describe the scene, actions, and filming approach for each shot, and produce editable short clips. For example, create shots of a character walking on grass or pausing on a street, first verify appearance consistency, then select results that meet narrative requirements to assemble into a series.
Creative Previsualization for Costumes and Props
Prepare character, costume, or prop assets, clearly specify the purpose of each reference image in the prompt, and generate dynamic previews of characters carrying props or presenting looks. Deliverables are suitable for discussing visual direction, action arrangements, and scene atmosphere; when product details are involved, check whether textures, logos, and structures match the original design.
Multi-Aspect-Ratio Content Assets
Using the same set of character reference assets, create landscape, portrait, or square short clips for creative tests in different display placements. When providing input, adjust both the aspect ratio and composition description rather than changing only the ratio; after output, check subject placement and action space before passing it to the editing workflow to add subtitles and brand information.
How to Choose This Model
Reference Images or a Fixed First Frame
If you want a character to continue appearing in a new environment without requiring the reference image to become the opening frame exactly as-is, choose happyhorse-1.1-r2v. If the task focuses on animating an approved image as the first frame, choose happyhorse-1.1-i2v; if there are no image assets and you mainly rely on script and shot text, use happyhorse-1.1-t2v.
Distinguish Versions and Editing Tasks
Both 1.1-r2v and 1.0-r2v generate from reference images. The stated difference in public specifications is that 1.1 adds a native 480P option, while this entry still provides 720P and 1080P. Do not interpret version numbers as a fixed quality increase. If you already have a video that needs costume changes, style transfer, or partial replacement, choose happyhorse-1.0-video-edit.
Get Started
Prepare Input for the Corresponding Operation
Prepare 1–9 ordered image_urls, and use character1, character2, and so on in the prompt to reference characters or subjects by their corresponding sequence numbers.
Explicitly Specify the Version
Set model=happyhorse-1.1-r2v and action=reference_to_video for /happyhorse/videos. Choose an integer duration of 3–15 seconds and 720P or 1080P; text and reference-image generation can set ratio.
Query and Save Completed Videos
Use async or callback_url to integrate background tasks, save task_id and query /happyhorse/tasks; wait for succeeded before reading video_url, then check the characters, actions, and audio.
Usage suggestion: a reference narrative with two characters
Input and objective
character1 and character2 are defined by the first two character reference images. They meet in the bookstore scene provided in the third image and exchange a book, with simple actions.
Acceptance criteria and next steps
Ensure the reference image order matches the character references, and choose an appropriate aspect ratio; check whether the characters are swapped, as well as hand movements and the background. Complex continuous stories still require storyboards.
Usage boundaries
Reference-image generation is neither video editing nor motion copying. The creative inputs for this model are images and prompts; do not treat video_url in a shared interface as its video reference capability. When you need to preserve existing camera movement or modify the original footage, use the appropriate editing workflow.
A single generation lasts 3–15 seconds, making it unsuitable for directly delivering long-form continuous narratives. You can first split it into independent shots and then edit them into a finished video; character details, spatial relationships, and action continuity across clips still require manual review, and reference images cannot replace a complete continuity review.
Images must be submitted as publicly accessible URLs, with their order consistent with prompt references. If multiple assets make conflicting requests for the same character's clothing, appearance, or style, the creative intent becomes unclear; it is recommended to curate the assets and clearly specify the role of each image.
Frequently Asked Questions
Will the reference image directly become the first frame of the video?
This model uses reference images to guide characters and visual elements, and should not be used as a fixed first-frame workflow. If you need an image with a specific composition as the opening frame, choose happyhorse-1.1-i2v. r2v is better suited for preserving the reference subject while rearranging the environment, actions, and camera through text.
How do I make prompts correspond to multiple reference images?
Put 1–9 image URLs in image_urls, and reference them in order in the prompt using character1, character2, and so on. In addition to specifying the corresponding assets, clearly describe the subject's actions, scene, and the purpose of each asset. Avoid merely listing names without explaining their relationships.
Which key items need to be explicitly specified when calling it?
Set action to reference_to_video, model to happyhorse-1.1-r2v, and submit prompt and image_urls. Do not rely on the default text-to-video action. Then choose resolution, ratio, and duration according to your needs to form a complete reference-image generation request.
Can it generate audio or preserve the original video's sound?
This model has native audio capabilities, but reference-image generation does not mean preserving the original video's sound, nor does it mean specifying voice-over or lip sync. The origin usage for preserving original audio belongs to video editing tasks; if audio has specific delivery requirements, check the generated result and arrange any necessary audio post-production.
How do I retrieve asynchronously generated videos?
After submitting with async, save the task_id and query the task through /happyhorse/tasks; you can also provide callback_url to receive completion notifications. The result includes the task status and video_url. Statuses may be pending, succeeded, or error. Download the video for inspection after successful completion.