All models

maestro

maestroVideo
Get your API key
maestro

From natural-language ideas to multilingual videos with subtitles

Maestro is a video production service centered on an AI director. After you describe the topic, audience, and communication goal, it organizes the script, visuals, voice-over, music, subtitles, and rendering to turn an idea into a complete video. It is suited for educational explainers, product promotion, and multilingual content creation, and can also incorporate image, video, and audio assets to continue editing or extending existing projects.

MaestroModel brand
VideoModel type
Script · Voice-over · SubtitlesCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Creation input
Natural-language prompt; image, video, and audio URLs may be attached, up to 20 items
Target duration
5–300 seconds, 30 seconds by default
Aspect ratio options
9:16, 16:9, 1:1; 9:16 by default
Multilingual production
Up to 4 languages per request; zh-cn by default, with the first as the primary language
Video types
auto、narrated、captions、avatar、drama
Creation actions
generate、remix、edit、extend
Delivery method
Asynchronous task; retrieve the completed video with subtitles and each language version after completion

The above are the creation and API specifications for the Maestro video production entry point. The target duration does not equal the exact final video duration.

Core capabilities

Organize ideas into complete videos

Maestro does more than generate visuals: it connects topic selection, scripts, narration, music, subtitles, and editing into a production workflow. Prompts can clearly specify who the audience is, what needs to be explained, how the opening should capture attention, and how the ending should conclude, keeping creation focused on the communication goal and making it suitable for starting production directly from a content brief.

One set of visuals, multilingual expression

Specify languages through langs, with the first serving as the primary language. Other languages reuse the same visuals before their corresponding voice-overs and rendering are completed. Chinese, English, Japanese, and other versions can be organized in the same task, reducing the work of repeatedly preparing visual content; delivery results are separated by language for convenient individual publishing.

Continue creating from existing projects

After a video is completed, there is no need to start from scratch every time. remix preserves the theme while adjusting presentation, edit is for local changes such as titles, voice-overs, or color schemes, and extend is for expanding content. Provide a historical task ID and clear modification requirements to start a new iteration task and progressively refine the same video project.

Use Cases

Educational Explainers and Tutorial Shorts

Enter a concept, tutorial key points, or organized article content, specify the audience's knowledge level and the conclusions you want to retain, then choose narrated to create an explainer video. Deliverables include visuals, narration, and subtitles, making it suitable for turning text content into easy-to-watch short videos; key terms and required information should be stated in the prompt.

Product Promotion and Brand Content

Attach reference materials such as product images and logos, describe the selling points, audience, and desired visual character, then choose a landscape, portrait, or square aspect ratio. Use presets such as modern and luxury to shape the overall look, and pair them with a narration voice to support the message; after the video is produced, continue editing the title or voice-over to create different promotional expressions.

Existing Video Processing and International Distribution

When adding subtitles to existing material, use captions and submit the source video URL; for cross-language distribution, select target languages in langs. The former focuses on processing existing videos, while the latter focuses on delivering the same visual content in multiple languages, and they can be used respectively for material organization and multi-region releases of product introductions.

How to Choose This Model

Choose Maestro When You Need a Complete Video

If the task requires not only visuals but also a script, voice-over, music, and subtitles, Maestro's director-style workflow is better aligned with a complete delivery goal. Rather than making requests only around a single segment of visuals, prioritize describing the content structure and audience expectations here. Use generate for new projects; for projects already created, choose the action based on whether you want to reinterpret, make local edits, or extend them.

Arrange Assets and Controls by Video Type

Choose narrated for explainer content, captions to subtitle existing videos, avatar for talking-head videos, and drama for character dialogue stories. The type determines the presentation format, style adjusts the visual feel, and voice adjusts the narration voice; do not treat all three as the same kind of control. When you want the system to organize the format itself, you can keep auto and clearly state the creative goal.

Get Started

Write a Finished-Video Brief

Specify the audience, topic, structure, duration, and style in prompt; provide assets through file_urls. captions requires a source video, and avatar requires a portrait.

Choose Video Type and Languages

Specify scenario, aspect, and target duration to /maestro/videos; langs supports up to 4 languages at a time. To modify an existing project, use remix/edit/extend with ref_task_id.

Review Every Language Version

Save task_id and query /maestro/tasks, then wait for final success; check the script, voice-over, subtitles, visuals, and volume separately, and do not treat task acceptance as completion of the finished video.

Trial suggestion: knowledge explainer video with subtitles

Input and objective

Create a 30-second explainer for beginners on “What is a vector database”: first explain its purpose, then use finding similar images as an example, and end with a memorable takeaway, with Chinese voice-over and subtitles, in a clean modern visual style.

Acceptance criteria and next steps

Use narrated or auto, focusing on verifying script facts, subtitles, voice-over, and asset transitions; up to four languages per request, and review each version after completion.

Usage limits

  • The target duration must be within 5–300 seconds, the maximum number of languages per request is 4, and up to 20 reference URLs are allowed. Multilingual production reuses the same set of visuals; this does not mean the visual content will be redesigned for every language version. If different regions require different visuals, separate creation tasks should be used.
  • captions requires a source video, and avatar requires a portrait; you cannot simply select a type while omitting the required assets. Reference content should be submitted through file_urls with media URLs; this differs from directly uploading local files, so accessible asset links should be prepared before production.
  • Maestro uses asynchronous production; successful submission only means the task has been accepted, not that the final video is complete. Major content adjustments may trigger regeneration; visual style and narration voice are creative controls and should not be regarded as tools for locking visuals frame by frame or guaranteeing voice-over results word by word.

Frequently Asked Questions

What is the difference between Maestro and tools that only generate video visuals?

Maestro is designed to organize complete video production, including not only visuals but also scripts, voiceovers, music, subtitles, editing, and rendering. Prompts should describe the content objective and presentation structure, not just the appearance of shots; it is better suited for creative tasks that need to progress from a brief to a finished video.

How do I use my own product images, videos, or audio?

Put media URLs in file_urls, and explain the purpose of each asset in the prompt, for example, product images for display, logos for brand identification, and videos for subtitle processing. You can submit up to 20 items; when selecting captions, provide a source video, and when selecting avatar, provide a portrait.

How many languages can be produced at once? Will the visuals differ?

You can specify up to 4 languages per request, with Chinese as the default, and the first language in langs is the primary language. The remaining languages reuse the same visuals, with separate dubbing and rendering, and language versions are provided upon completion. If you want different visual narratives for different regions, submit separate creative requirements for each.

Should I choose remix, edit, or extend to modify a finished video?

Choose remix if you want to keep the theme but use a different presentation; choose edit to change specific elements such as the title, voiceover, or color scheme; choose extend to continue expanding the original content. All three require ref_task_id, and you should describe the modification goal in the prompt; a new task ID will then be returned.

How do I get the final video after submission?

After calling POST /maestro/videos, first obtain task_id, then use POST /maestro/tasks to check progress and results. The task goes through planning and production stages, and upon success you can obtain the finished video and each language version; applications should distinguish between submitted successfully, in production, completed, and failed states.