All models

grok-imagine-video

xAIVideo
Get your API key
grok-imagine-video

Conceptualize shots with text and turn static images into dynamic videos

grok-imagine-video is the video generation entry point for the xAI Grok Imagine series, suitable for turning text ideas or existing images into dynamic clips. It focuses on text-to-video and image-to-video: the former starts with a scene description, while the latter uses an image as a visual starting point and adds motion intent. Through this platform, you can submit generation tasks, track their status, and obtain video links, making it suitable for short-video assets and creative previsualization.

xAIModel brand
VideoModel type
Text · Image guidanceCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
grok-imagine-video
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Creation method
Text-to-video, image-to-video
Primary inputs
prompt text; image_url image link
Video controls
The API provides aspect ratio, resolution, and duration settings
API endpoint
POST /grok/videos; model specifies grok-imagine-video
Result delivery
JSON task results and video_url video links
Task management
Asynchronous submission, status queries, completion callbacks

The creation method and task management descriptions correspond to this platform's API endpoint; aspect ratio, duration, and resolution settings are not equivalent to the model's native limits.

Core capabilities

Directly conceptualize dynamic shots from text

When no image is available, you can use prompts to describe the subject, environment, action, and camera intent to start text-to-video creation. This is suitable for first exploring whether a visual concept works before deciding whether to move into filming or post-production. Prompts should focus on clear actions and avoid packing multiple unrelated scenes into the same description.

Give existing images motion expression

When you have product images, character illustrations, or scene images, you can submit an image through image_url and then describe the desired action in text, such as the subject turning around, the camera slowly pushing in, or changes occurring in the environment. The image provides a visual starting point, while the prompt adds motion direction, making it suitable for animating existing design assets.

Integrate video generation into task workflows

Generation tasks can be submitted asynchronously, without requiring business pages to continuously wait for complete results. After saving the task_id, an application can query progress or use a completion callback to receive results, then display the video link based on the status. This workflow is suitable for asset workbenches, creative review pages, and automated content production programs.

Applicable Scenarios

Short Video Creative Previsualization

Enter a shot concept, specifying the subject, action, and atmosphere in the frame, to generate a video draft for discussion. Teams can further refine options around motion direction, compositional intent, and pacing choices. The deliverable is dynamic previsualization material, not a fully finished video with subtitles, music, and editing already completed.

Animating Static Product Assets

Use existing product images as input, add descriptions of display actions or camera changes, and create candidate dynamic assets. Suitable for exploring ways to present product displays, campaign visuals, or brand content. Before formal release, check each item to ensure product outlines, logos, and details still meet the original design requirements.

Motion Exploration for Illustrations and Scenes

Submit character illustrations or environment images, and use action prompts to describe changes in facial expressions, poses, or scene movement, obtaining clips for creative discussion. Adjust descriptions around the same input, compare different motion approaches, then select suitable results for subsequent editing, rather than expecting complex narratives to be completed in one pass.

How to Choose This Model

Choose It When You Need Both Text and Image-Based Creation

If a project includes both text-only concepts and existing images that need animation, the two creation modes of grok-imagine-video make it easier to organize a unified workflow. Compared with grok-imagine-video-1.5:official, the latter explicitly requires image input and is not suitable as a direct replacement for text-only ideation. When choosing, first determine the starting point of your assets rather than looking only at the version name.

Distinguish Related Versions by Specific Endpoint

grok-imagine-video, public invocation IDs with :official or :reverse, and variants containing 1.5 should be managed as separate configurations and should not be treated as completely interchangeable names. When reusing an existing workflow, explicitly specify model; when migrating to a related version, retest input modes and parameter combinations, rather than directly applying new-version settings to this model.

Getting Started

Choose a Text or Image Starting Point

For text-to-video, provide a prompt; for image-to-video, use image_url to define the visual starting point and add action descriptions. Configure reference image capabilities according to the guide for this public variant.

Specify the Full Public Invocation ID

Specify model=grok-imagine-video for /grok/videos, starting with a small 6-second task; use the current parameter range in the API panel for duration, and do not directly apply duration specifications from other variants. Choose aspect_ratio and resolution according to the actual visuals.

Retrieve Generation Task Results

Use async or callback_url to track the task, and use task_id to query /grok/tasks; wait for succeeded before reading video_url, check the subject, motion, visual continuity, and ending, and proceed to editing with the video actually returned.

Trial suggestions: single-shot exploration with text and images

Input and objective

A wooden toy train moves slowly along a tabletop, with the camera following at a low angle, maintaining the train's appearance and a simple background.

Acceptance and next steps

Explicitly enter the complete ID without a suffix, and set parameters according to the API panel; do not directly treat the duration or resolution of another 1.5 variant as guaranteed for this ID.

Usage boundaries

  • This model should be used for video clip generation; do not treat it as a Grok conversational model or full editing software. Text Q&A, web search, video extension, audio track creation, and multi-shot editing should be arranged as separate requirements, and must not be assumed to be included in a single generation merely because they belong to the Grok family.
  • Aspect ratio, resolution, and duration are configuration dimensions of the video interface, but the available combinations differ across models. Save an independent generation configuration for this model, verify the target aspect ratio and output settings before batch production, and do not directly apply the highest specifications of related 1.5 variants.
  • Image inputs are used to guide video creation and should not be treated as editing instructions that lock the original image frame by frame. Character appearance, product text, and fine patterns should be manually checked in the results; when continuous storytelling is needed, first split it into independent shots, then organize the sequence and pacing through editing.

Frequently Asked Questions

Can I use grok-imagine-video without an image?

You can start creating text-to-video content from a text prompt. Simply describe the subject, environment, and actions you want to appear; there is no need to prepare a static image first. If you already have a clear visual design, you can instead use image input and focus the prompt on motion and camera changes.

What should I write in the prompt for image-to-video?

Prioritize describing what happens next in the image, such as a person turning around, an object moving, or the camera pushing in, rather than merely repeating what is already in the image. After submitting an image through image_url, action prompts can help convey creative intent; the result should still be checked for subject appearance and visual details.

What is the difference between it and grok-imagine-video-1.5:official?

They are different invocation IDs. grok-imagine-video is intended for creative workflows starting from text and images, while grok-imagine-video-1.5:official explicitly uses image-to-video and requires image_url to be submitted. When migrating, choose again based on the input material rather than simply replacing the name.

Can I directly select fun, normal, or spicy?

Do not submit these names as dedicated switches for this model. When creating, naturally describe the desired atmosphere, actions, and presentation in the prompt, such as relaxed, restrained, or dramatic; style descriptions express generation intent and are not deterministic effect presets.

How do I obtain the video after asynchronous submission?

Save the task_id returned by the submission, track the status through task queries, or set callback_url to receive completion notifications. When the state in the result is succeeded and video_url exists, then display or retrieve the video; pending means it is still being processed, while failed should enter the error-handling flow.