All models

grok-imagine-video-1.5-fast:reverse

xAIVideo
Get your API key
grok-imagine-video-1.5-fast:reverse

Text and Image-to-Video Generation for Short-Form Creative Iteration

grok-imagine-video-1.5-fast:reverse is the video generation entry point in the xAI Grok Imagine series, designed to turn text concepts, product photos, and storyboard frames into short videos. It supports text-to-video, image-to-video, and reference image guidance; you can set the aspect ratio, resolution, and duration, with video links delivered through asynchronous tasks, making it suitable for creative test shoots and asset production.

xAIModel Brand
VideoModel Type
Text · Image GuidanceCreation Method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
grok-imagine-video-1.5-fast:reverse
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API Features

Creation Method
Text-to-video, image-to-video, reference image guidance
Input Assets
prompt text, image_url image link, reference_image_urls reference image list
Generation Duration
6–30 seconds, 6 seconds by default; recommended to start with 6 or 10 seconds
Resolution Settings
480p, 720p, 1080p; 480p by default
Aspect Ratio Settings
1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3
Result Delivery
Asynchronous task ID, task status, and video_url video link
API Endpoint
POST /grok/videos, with model set to grok-imagine-video-1.5-fast:reverse

The durations and settings above describe how this platform invokes the model and do not represent the native model's specifications in every usage environment.

Core Capabilities

Start Test Shoots from Text Concepts

When no existing images are available, you can directly describe the subject, environment, actions, and camera movement to generate a video. Text-to-video requires a prompt and is ideal for first exploring visual directions, then refining actions and composition iteratively; clearly writing the main intent of a single shot makes it easier to compare results than cramming an entire complex story into the same prompt.

Bring Static Assets into Dynamic Expression

Image-to-video receives images through image_url and can be paired with prompts describing the desired motion. When you already have product photos, character visuals, or storyboard sketches, there is no need to rebuild the visual starting point from text alone; a reference image list can also guide style and content, but reference guidance should not be interpreted as strictly locking every detail.

Manage Video Delivery with Task Workflows

Generation tasks can be submitted asynchronously through async or callback_url, so application pages do not need to keep waiting for a connection. After saving the task_id, you can check progress or receive callbacks, and obtain the video_url when the status is succeeded; this approach is suitable for separating generation, review, and asset saving into clear processing steps.

Use Cases

Animating Product Photos

Upload a product image and describe the display action, camera angle, and background atmosphere to generate dynamic clips for creative review. You can try different shot descriptions around the same asset, deliver multiple video candidates, and then have the team check whether the product appearance, motion effects, and visuals meet promotional requirements.

Social Short-Form Video Concepts

Write the campaign theme or scene concept as a prompt, choose a vertical, square, or landscape aspect ratio suitable for the publishing placement, and create short video candidates. It is recommended to first define the subject's actions and camera rhythm, then adjust the background and visual style; the delivered video assets can continue into editing, subtitles, and brand packaging workflows.

Storyboards and Shot Previsualization

Use storyboard frames as input images, together with descriptions of character actions and camera movement, to observe how static compositions perform once in motion. This is suitable for discussing shot direction before formal filming, or comparing different dynamic approaches for the same frame; results support creative judgment rather than replace professional production with frame-by-frame control.

How to Choose This Model

When You Need to Start with Text or Create Longer Clips

If you have not yet prepared a source image, or want to generate a clip longer than 15 seconds at once, consider this model first: it supports starting from text alone, and the usage guide gives a duration range of 6–30 seconds. By comparison, grok-imagine-video:reverse has a range of 1–15 seconds. Your choice should be based on input assets and clip length, rather than judging speed from the name alone.

How to Choose Between This and the 1.5 Image-to-Video Version

grok-imagine-video-1.5:official requires image_url and is suitable for tasks that clearly use an image as the creative starting point; this model retains both text-to-video and image-to-video options, making it convenient for use at different stages of asset preparation. They are different invocation IDs; when switching, select the duration and input combination again, and do not directly apply one version's workflow to the other.

Get Started

Choose a Text or Image Starting Point

For text-to-video, provide a prompt; for image-to-video, use image_url to define the visual starting point and add action descriptions. Set reference image capabilities according to this publicly available variant guide.

Specify the Full Public Invocation ID

Send model=grok-imagine-video-1.5-fast:reverse to /grok/videos, starting with a small 6-second task; durations range from 6–30 seconds. Choose aspect_ratio and resolution based on the actual visuals.

Retrieve Generation Task Results

Use async or callback_url to track the task, and use task_id to query /grok/tasks; wait for succeeded before reading video_url, check the subject, actions, visual continuity, and ending, then move the returned video into editing.

Trial suggestion: a draft for longer actions

Input and objective

The person in the reference image walks into the room from beside the door, puts down their backpack, and sits at the desk. The camera follows slightly, maintaining the clothing and scene.

Acceptance and next steps

Verify the action sequence with a 10-second task; for longer durations, set it within the 6–30 second range. Record the full suffix separately from other Grok variants.

Usage boundaries

  • This model is more reliable when tasks are planned for 6–30 seconds; do not treat it as a video creation tool for arbitrary lengths. The guide recommends starting with 6 or 10 seconds; complex stories can first be split into separate shots, then organized during editing rather than relying on a single request to complete the entire narrative.
  • Reference images are used to guide style or content; they do not represent first-and-last-frame control or frame-by-frame reproduction. Character appearance, product details, and action continuity should be checked item by item in the finished video; projects requiring strict visual consistency should retain human review and subsequent refinement steps.
  • This introduction does not list audio, lip-syncing, video editing, or extending existing videos as capabilities of this model. Successful task submission does not mean the video is complete; process results according to the pending, succeeded, and failed statuses, and obtain the video link after success.

Frequently Asked Questions

Can it generate without uploading an image?

Yes. This model supports text-to-video; simply provide a prompt describing the scene, without needing to submit an image_url at the same time. The prompt can specify the subject, actions, environment, and camera movement; if you already have a clear product or character image, you can instead use image-to-video to establish a visual starting point.

Do I still need a prompt after providing an input image?

After submitting image_url for image-to-video, the prompt is optional, but it is recommended for describing the desired actions and camera changes. The image provides the visual material, while text supplements the motion intent; for example, explain how the subject should move and how the camera should approach, rather than merely repeating content already present in the image.

What is the difference between reference images and input images?

image_url is used as the input image for image-to-video, while reference_image_urls are used to guide style or content. Reference images can help express the creative direction, but should not be understood as mandatory duplication or specified first and last frames; if the goal is to preserve the visual starting point of an image, prioritize image-to-video.

How should I choose duration and resolution?

This model uses a duration range of 6–30 seconds, with 6 seconds as the default; the guide recommends trying 6 or 10 seconds first. Resolution settings provide 480p, 720p, and 1080p, with 480p as the default; you can first verify the motion and composition, then choose other settings based on the intended use and check the actual result of the final video.

How do I know when the video has finished generating?

After asynchronous submission, save the task_id and query the task through POST /grok/tasks, or use callback_url to receive a completion notification. pending means it is still being processed, succeeded means success, and failed means failure; after success, read video_url in data to obtain the video, and do not rely solely on the submission response to determine whether the finished video is ready.