Native Sound Short-Film Model Integrating Text, Image, Audio, and Video References
MiniMax-H3 is a general-purpose multimodal audio-video model for short-film creation. It can combine text, images, video, and audio to understand creative intent and generate videos with native stereo sound. It is suitable both for building shots from scripts and for controlling transitions with first and last frames, or for guiding characters, products, actions, and sound with various reference materials. It is ideal for advertising, character narratives, and content with audiovisual rhythm.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Video duration
4–15 seconds, whole-second durations
Output resolution
768P, 2K
Aspect ratio control
21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and adaptive
Native audiovisual specifications
24 FPS video; 32 kHz stereo audio
Creation methods
Text-to-video, first frame/last frame/first and last frames, multimodal reference generation
Number of references
Up to 9 images, 3 video clips, and 3 audio clips; up to 12 files in total
Text and delivery
Text entries up to 7,000 characters; supports waiting for results, asynchronous queries, and callbacks
Frame rate and audio sampling rate are native model specifications. This platform entry provides controls for resolution, duration, aspect ratio, and asset roles.
Core Capabilities
Create visuals and sound together
H3 jointly generates video and stereo sound, so visual concepts and sound design do not need to be completely separated. Prompts can simultaneously describe subject actions, camera changes, ambient sound, dialogue, and musical atmosphere, allowing sound to serve scene expression rather than simply adding a background track after the video is finished.
Use first and last frames to constrain shot direction
Providing a first frame can bring existing product images, posters, or character visuals to life; providing a last frame can specify the final camera destination; providing both is used to generate the transition in between. The creative focus is to describe the transition action and keep the composition, proportions, and lighting of the two images as consistent as possible.
Give different assets distinct roles
Reference images guide characters, clothing, products, or style; reference videos guide actions and camera movement; reference audio guides timbre, music, or rhythm. Clearly defining the role of each asset is better suited to organizing complex creations and makes it easier to determine whether the final video retains key characteristics.
Use Cases
Product and Brand Shorts
Provide product images, model images, and scene references; describe product close-ups, how people use the product, and the ending presentation to generate landscape or portrait advertising assets. Specify the person's appearance and the product structure separately, and before delivery, focus on checking contours, wearing relationships, reflections, and brand text.
Character Narratives and Storyboard Previsualization
Provide character and space references, along with descriptions of character relationships, emotional changes, shot sizes, and dialogue, to create short drama clips or storyboard previsualizations. Use close-ups to convey gaze and expressions, use reference images to guide clothing and appearance, then incorporate suitable short clips into subsequent editing workflows.
Motion and Music Rhythm Content
Use action videos, character images, and audio as references; specify who performs the action, which shots to use, and how sound should complement the visuals to generate dance or fashion rhythm shorts. First trim assets into short segments relevant to the target clip to prevent unrelated content from distracting from the creative focus.
How to Choose This Model
Choosing Between H3 and H3 Max
Choose H3 when you need 2K output, short shots starting at 4 seconds, or complete audiovisual creation based on multimodal assets. H3 Max is another speed-optimized model with native output at 480P or 768P and durations of 5–15 seconds; consider it when fast generation is more important, but do not mix up the specifications of the two.
Choose a Creation Method by Control Goal
When no ready-made assets are available, use text to clearly specify the subject, action, camera, and sound; when you already have defined opening or ending frames, choose first-and-last-frame mode; when you need to combine character appearance, actions, and voice timbre, choose reference mode. First-and-last frames and reference assets cannot be included in the same request, so determine the most important control goal first.
Getting Started
Use content to Express Shots and References
Include at least one non-empty type=text entry; set images, videos, and audio by role. First-and-last-frame scenarios and reference asset scenarios are mutually exclusive, so do not submit both sets of controls at the same time.
Select the Output Range for This Model
Set model=MiniMax-H3 for /minimax/videos, and choose 768P/2K and an integer duration of 4–15 seconds. Explicitly select an aspect ratio for text and reference modes; first-and-last-frame mode uses adaptive based on the input images.
Receive Task content.url
After asynchronously saving task_id, query /minimax/tasks or receive a callback, wait until task.status=succeeded, then read task.content.url; check the audio, visuals, and duration, and download it for editing.
Trial Recommendations: Multimodal Rhythm and Character References
Inputs and Goals
A character image defines the appearance, video defines simple dance moves, and audio defines the rhythm; the character performs a sequence of movements on a brightly lit stage, retaining the clothing and main subject, with a steady camera.
Acceptance and Next Steps
Distinguish assets by reference_image/reference_video/reference_audio, and avoid mixing them with first_frame/last_frame; check the movements and native sound.
Usage Limits
Each generation is limited to 4–15 seconds, with a resolution of 768P or 2K. Longer narratives need to be split into shots and then edited; character, space, and sound continuity across segments should be checked segment by segment, and a single short-video generation should not be treated as long-video production.
Each reference video and audio clip must be 2–15 seconds, with the total duration of each type not exceeding 15 seconds; mixed references may include up to 12 files in total. More reference materials are not necessarily better; clearly specify what to retain, and avoid contradictory movement, styling, or sound requirements.
Asset references and first/last-frame controls are used to guide generation and do not guarantee pixel-perfect reproduction. Product details, text, complex movements, and dialogue transitions still need to be checked; when involving character likenesses, voice timbres, and brand assets, use materials with legal authorization.
Frequently Asked Questions
Can MiniMax-H3 generate videos using only text?
Yes. Provide a non-empty text item in content to start text-to-video generation, while also setting resolution, duration, and a fixed ratio. Prompts are recommended to describe the subject, action, scene, camera, lighting, and sound in that order; adaptive aspect ratios cannot be used in text-only mode.
What is the difference between a first-frame image and a reference image?
first_frame is used to specify the opening frame of the video, while reference_image is used to guide the subject's appearance, product, or style and does not need to become the opening. Use last_frame to control the final frame; first-and-last-frame mode cannot be used together with reference_* assets.
Can audio be used to guide sound and rhythm?
reference_audio can be used in reference mode, and the text should specify whether it is for timbre, dialogue, music, or rhythm. Audio assets support WAV and MP3; H3 natively supports stereo generation, and both Chinese and English are stably supported dialogue languages.
Is 2K output just ordinary upscaling?
H3's native 2K workflow uses regeneration, combining the base video with the original context to restore details rather than relying only on a traditional super-resolution module. Set resolution to 2K when calling it to specify the output requirement; there is no need to organize these internal generation steps yourself.
How do I submit a task and retrieve the finished video?
Submit MiniMax-H3 and content assets to /minimax/videos; it waits for completion and returns a task by default. Pass async=true or callback_url to immediately obtain a task identifier, then query it through /minimax/tasks or receive a callback; after success, retrieve the finished video from task.content.url.