Can OmniHuman 1.5 generate a talking video from text alone?
This workflow requires a portrait image and driving audio; text alone cannot replace these two assets. You can first record or create the script as an mp3/wav, then submit the audio URL. The prompt is used to adjust expressions, emotions, and style; it does not directly read the lines aloud.
What requirements should the photo meet?
It is recommended to use a clear, front-facing portrait with good lighting and no obstructions. The person should occupy an appropriate portion of the image. The photo must have a publicly accessible URL. Before formal production, you can test the photo with a short audio clip to check lip movements, expressions, and head-and-shoulder motion, then continue using the asset.
Can I make a specific person in a group photo speak?
You can submit an array of subject mask URLs through mask_url to specify the person to drive in a multi-person photo; even if there is only one mask URL, it should still be submitted as an array. You should first identify the target and prepare the corresponding mask. This feature is for subject selection and should not be understood as allowing multiple characters to speak separately or complete a multi-person dialogue in a single request.
How can I control the tone and character movements?
The voice itself provides speaking speed, pauses, and delivery rhythm, while the prompt can add requirements such as a gentle, calm, natural manner or slight head movements. It is recommended to keep the voice-over and prompt consistent, avoiding an excited voice with text requesting calmness; the final performance should still be checked through the generated video.
How do I submit and retrieve videos in the application?
Submit image_url and audio_url to POST /dreamina/videos, and specify omnihuman-1.5. Results can be retrieved synchronously for short tasks; for longer tasks, you can use callback_url, or set async:true and then query /dreamina/tasks. After completion, read video_url and save it.