All models

omnihuman-1.5

DreaminaVideo
Get your API key
omnihuman-1.5

Generate natural lip-syncing digital humans from a single portrait and voice

OmniHuman 1.5 is an audio-driven video model for talking-head content: starting with a single portrait photo and a voice clip, it makes a static image speak while coordinating lip movements, expressions, and head-and-shoulder motion. It is suitable for product explainers, course introductions, and brand announcements. You can adjust emotion and style with prompts, and use subject masks to specify the on-camera person in group photos.

DreaminaModel brand
VideoModel type
Audio-driven portraitCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
omnihuman-1.5
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Creation method
Generate talking-head videos from a single portrait image + driving audio
Input assets
Publicly accessible image URLs and mp3/wav audio URLs
Audio recommendation
Recommended to keep under 60 seconds; not a hard duration limit
Performance control
Use prompts to adjust expressions, emotion, stability, and style
Subject selection
mask_url supports selecting the driven person in group photos
Tasks and output
Synchronous or asynchronous generation, returning task status and video_url

The above are this platform's asset requirements and invocation methods; the recommended audio duration does not represent the model's native limit.

Core capabilities

Voice drives complete talking-head performance

OmniHuman 1.5 does more than animate the mouth in a photo: it coordinates expressions and head-and-shoulder movements with the rhythm of the voice. The creative focus is on how the person delivers the content, making it suitable for talking-head videos that require a consistent presenter; when preparing audio, first determine pauses, emphasis, and the overall delivery tone.

Use prompts to shape the delivery tone

In addition to the photo and voice, you can use prompts to describe emotions such as gentleness, calmness, and naturalness, and request subtle head movements or stable performance. This guides the person's performance style, making it suitable for creating course explanations, brand announcements, and warm content separately, rather than replacing the dialogue in the audio.

From portrait assets to downloadable video

Use a clear front-facing portrait as the visual foundation, without preparing multi-angle modeling assets in advance. Group photos can be paired with subject masks to specify the driven subject; after generation, receive a video URL to continue downloading, editing, or adding it to product pages, allowing talking-head content to fit into existing content production workflows.

Applicable Scenarios

Product Updates and Feature Explanations

Turn product updates into voiceovers, pair them with a fixed presenter headshot, and generate talking-head clips introducing new features. Then combine them with screen recordings, subtitles, and key annotations to deliver videos for help centers or new-user onboarding, with the person handling narration and the screen recording demonstrating actual operations.

Course Introductions and Training Materials

Record course openings, knowledge-point overviews, or training reminders as audio, pair them with an instructor photo, and generate chapter-based digital human explanations. Using the same avatar maintains visual continuity across the course; when formulas, charts, or steps are involved, add courseware visuals in post-production instead of making the person present all the information.

Brand Talking Heads and Content Reuse

Prepare different audio for event announcements, product introductions, or frequently asked questions, and use the same branded persona to generate separate videos. Deliverables can continue to include product footage, subtitles, and end cards; reusing the photo can reduce repeated filming, but each piece of content should still be checked to ensure expressions, lip movements, and audio match.

How to Choose This Model

Choose It When Photos and Voiceovers Are Ready

If you already have a usable portrait photo and finalized voiceover, and the goal is to have the person speak directly to the audience, OmniHuman 1.5's single-image audio workflow is straightforward. It is better suited to performances organized around sound rather than building complex narratives from text. Choose 1.5 based on talking-head needs and sample results; there is no need to interpret the version number as a comprehensive improvement for every task.

Distinguish Talking Heads from Complex Scene Animation

For presenter explanations, course hosting, and brand announcements, prioritize OmniHuman 1.5; if the focus is on large movements, character positioning, scene changes, or precisely recreating an action sequence, consider animation or video models that support the relevant controls. For talking-head tasks, first test photos, voices, and prompts with short clips before producing the full content.

Get Started

Complete the Script and Audio First

Choose a clear portrait, prepare MP3/WAV audio, and verify the script. Both the image and audio must be publicly accessible; it is recommended to test with a short talking-head segment first.

Submit an Audio-Driven Portrait Task

Provide model=omnihuman-1.5, image_url, and audio_url to /dreamina/videos; prompt can describe emotion and stability, and for multi-person scenes provide mask_url according to the documentation.

Check Lip Sync and Expressions

Save the task_id asynchronously, query /dreamina/tasks or receive a callback; retrieve video_url, compare it against the audio to check mouth movements, pauses, and head-and-shoulder motion, then add subtitles.

Trial recommendation: audio-driven course introduction

Input and objective

Use a clear single-person portrait and a pre-recorded course introduction voiceover. The person speaks naturally, with a calm expression and slight head and shoulder movement, while the background remains stable.

Acceptance and next steps

Listen to the audio first, then check lip sync, pauses, and expressions; both image and audio URLs must be provided, and the prompt cannot replace the audio.

Usage boundaries

  • Portrait materials affect lip sync and expression quality. Prioritize photos with even lighting, a clear front-facing view, an unobstructed face, and an appropriate subject size; profile, blurry, or overly dark images are not suitable for direct use as final production materials. For group images, the target person must be clearly specified; specifying the subject should not be understood as multiple people speaking at the same time.
  • Driving audio must be prepared in advance; a text script cannot replace the required audio material. Both image and audio addresses must be publicly accessible, and the audio should be mp3/wav, preferably kept within 60 seconds. Longer content can be split by chapter for production, making it easier to check lip sync and performance effects segment by segment before post-production stitching.
  • Prompts are used to guide emotion and style; they are not tools for frame-by-frame action choreography, nor do they guarantee identical performances every time. After output, check lip sync, character movements, and content delivery, and save the video promptly; when using the likeness and voice of real people, obtain the appropriate authorization first.

Frequently Asked Questions

Can OmniHuman 1.5 generate a talking video from text alone?

This workflow requires a portrait image and driving audio; text alone cannot replace these two assets. You can first record or create the script as an mp3/wav, then submit the audio URL. The prompt is used to adjust expressions, emotions, and style; it does not directly read the lines aloud.

What requirements should the photo meet?

It is recommended to use a clear, front-facing portrait with good lighting and no obstructions. The person should occupy an appropriate portion of the image. The photo must have a publicly accessible URL. Before formal production, you can test the photo with a short audio clip to check lip movements, expressions, and head-and-shoulder motion, then continue using the asset.

Can I make a specific person in a group photo speak?

You can submit an array of subject mask URLs through mask_url to specify the person to drive in a multi-person photo; even if there is only one mask URL, it should still be submitted as an array. You should first identify the target and prepare the corresponding mask. This feature is for subject selection and should not be understood as allowing multiple characters to speak separately or complete a multi-person dialogue in a single request.

How can I control the tone and character movements?

The voice itself provides speaking speed, pauses, and delivery rhythm, while the prompt can add requirements such as a gentle, calm, natural manner or slight head movements. It is recommended to keep the voice-over and prompt consistent, avoiding an excited voice with text requesting calmness; the final performance should still be checked through the generated video.

How do I submit and retrieve videos in the application?

Submit image_url and audio_url to POST /dreamina/videos, and specify omnihuman-1.5. Results can be retrieved synchronously for short tasks; for longer tasks, you can use callback_url, or set async:true and then query /dreamina/tasks. After completion, read video_url and save it.