All models

happyhorse-1.1-i2v

HappyHorseVideo
Get your API key
happyhorse-1.1-i2v

Turn a single first-frame image into the starting point for a dynamic short video with sound

happyhorse-1.1-i2v is the first-frame image-to-video model in HappyHorse 1.1. It uses one image to define the starting point of the shot, then uses text to describe subject actions, environmental changes, and camera movement. It is suitable for turning product images, character design images, and scene illustrations into dynamic short videos, supports 3–15-second creation, includes native audio capabilities, and is especially suited to tasks with an established composition that need to further develop visual motion.

HappyHorseModel brand
VideoModel type
First-frame image-to-videoCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
happyhorse-1.1-i2v
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Creation method
Single-image first-frame image-to-video, with optional text action descriptions
Platform resolutions
720P, 1080P
Video duration
3–15 seconds; the platform accepts whole seconds
Native video output
24 fps, MP4
Aspect ratio control
The output aspect ratio follows the first-frame image as closely as possible
Native audio
Supports audio-video generation
Task delivery
Supports asynchronous queries, completion callbacks, and video URL results

Native resolutions also include 480P; this platform entry provides 720P and 1080P, and first-frame tasks determine the aspect ratio based on the input image.

Core capabilities

Set the visual first, then design the motion

The first-frame image directly serves as the starting point of the video, providing a clear basis for product placement, character appearance, and scene composition. Compared with building a shot entirely from text, this approach is better suited to creations with existing visual assets; prompts can focus on changes such as looking up, turning around, wind moving clothing, or a slow push-in, reducing repeated descriptions of the initial image.

Develop static assets into short shots

This model is designed for first-frame animation and can create short videos around subject actions, environmental dynamics, and camera movement. In practice, it is recommended to first define one primary action, then add lighting, background, and camera-motion requirements, allowing a few seconds of footage to serve a clear purpose rather than cramming a complete story, multiple transitions, and complex interactions into one clip.

Retrieve video results by task

Submit the first frame and creation description through POST /happyhorse/videos, explicitly selecting image_to_video and happyhorse-1.1-i2v. When background processing is needed, you can use async to obtain a task ID and query its status, or use callback_url to receive completion results, making it suitable for integration with asset production tools and batch creation workflows.

Applicable Scenarios

Animating Static Product Images

Input a clearly composed product photo, describe a subtle camera push-in, changes in background lighting and shadows, or movement in the surrounding environment, and generate a short clip for marketing previews. When packaging and brand text need to be highlighted, use restrained motion to preserve visual focus, and check labels, small text, and outlines before delivery to avoid dynamic changes affecting product information.

Character Shot Previsualization

Use a character design image as the first frame, combined with action descriptions such as looking up, looking back, and clothing movement, to obtain a dynamic storyboard draft. The deliverable can be used to discuss performance direction and shot pacing; if the task becomes one where multiple character images jointly constrain a new scene, choose a reference-image-to-video model instead of continuing to add first-frame inputs.

Dynamic Preview for Scene Illustrations

Start with a landscape image or scene illustration, describe the movement of water, leaves, clouds, and the camera, and create a short atmospheric shot. When producing vertical content, first prepare a vertically composed image so the aspect ratio continues from the first frame; after completion, obtain the video URL, download the assets, and enter the editing workflow to combine them into longer content.

How to Choose This Model

Choose a 1.1 Model Based on Input Assets

If you already have an image that needs to serve as the opening frame, choose happyhorse-1.1-i2v; if you only have a text script, choose happyhorse-1.1-t2v; if multiple images are needed to constrain characters, props, or style, choose happyhorse-1.1-r2v. The difference is whether the image must become the first frame, rather than merely serving as creative reference. Modifying an existing video belongs to a video-edit task.

Trade-offs with 1.0 and Other Video Tasks

The native duration, frame rate, and file format of the 1.1 and 1.0 first-frame models are the same; 1.1 adds a native 480P option. At this entry point, both use 720P or 1080P, so do not assume speed or image-quality improvements based on the version number alone. For new first-frame tasks, 1.1 is available; if you must specify the ending frame or connect long shots, choose a model that explicitly supports first-and-last-frame control.

Getting Started

Prepare Inputs for the Corresponding Operation

Prepare a single first-frame image_url, and use a prompt to describe the action; keep the aspect ratio aligned with the input image whenever possible, with no need to pass ratio.

Explicitly Specify the Version

Set model=happyhorse-1.1-i2v and action=image_to_video for /happyhorse/videos. Choose an integer duration of 3–15 seconds and 720P or 1080P, using the input image to determine the starting composition.

Query and Save the Completed Video

Use async or callback_url to integrate background tasks, save the task_id and query /happyhorse/tasks; wait for succeeded before reading video_url, then check the characters, actions, and audio.

Trial Recommendations: Portrait First Frame and Natural Motion

Input and Goal

Use a portrait photo as the first frame, with the person gently turning their head toward the camera while the background remains stable and their hair and clothing move naturally in a light breeze.

Acceptance and Next Steps

Set an integer duration of 3–15 seconds, and check the face, head, and audio; the platform only provides 720P/1080P, so do not claim that the platform also offers 480P just because it appears in the native table.

Usage Boundaries

  • This is a single-first-frame generation model, not a multi-image reference or existing video editing tool. Do not use image_urls in place of image_url, and do not expect to modify a video by passing video_url; the task action and model must match, and image_to_video should be explicitly specified when submitting.
  • Keep the aspect ratio aligned with the first frame whenever possible; it is not suitable for forcibly changing the composition by relying on ratio. Materials must use publicly accessible image URLs; it is recommended to complete landscape/portrait composition, subject spacing, and key content layout before uploading to avoid losing important parts of the image when cropping after generation.
  • The scope of a single creation is 3–15 seconds. The first frame defines the starting point; it does not mean that every subsequent frame will retain unchanged details. Complex occlusions, large turns, and fine text should be checked segment by segment. Audio capability also does not mean that voice-over, reference audio tracks, or precise lip-sync can be specified directly.

Frequently Asked Questions

How must the first-frame image be submitted?

Use image_url to submit a publicly accessible image, and set action to image_to_video and model to happyhorse-1.1-i2v. The image will serve as the first video frame; prompt is used to supplement subsequent actions, environmental changes, and camera movement, so there is no need to redescribe all static details.

Can I specify a vertical or square video?

The output aspect ratio for first-frame tasks should follow the image as closely as possible, with no need to pass ratio separately. To create vertical or square content, prepare a first frame with the corresponding composition first; it is best to leave safe space around product labels, people's heads, and other important elements, then check the actual final video frame.

Can it use multiple reference images to keep characters consistent?

This model starts generation from a single first-frame image and is not a multi-image reference mode. If you need to combine references for characters, clothing, or props, choose happyhorse-1.1-r2v; if the main goal is to make an existing image start moving without reorganizing the scene, i2v better matches the task structure.

What resolutions, durations, and audio capabilities are supported?

This endpoint offers 720P or 1080P, with durations of whole numbers from 3–15 seconds; native video specifications are 24 fps, MP4, with audio support. Audio generation and precise dubbing control are different capabilities; retaining the original video audio or using reference audio input cannot be used with this model.

How do I get the finished video after submission?

You can set async=true, save the returned task_id, and then query through /happyhorse/tasks; you can also provide callback_url to wait for a completion notification. Result statuses include pending, succeeded, and error, and successful results include video_url; do not assume the video has been generated just because you have obtained a task ID.