Image generation and reference image editing model balancing quality and speed
qwen-image-3.0 is Alibaba's standard image generation and editing model. It can create images from text descriptions and make modifications using reference images. It balances quality and speed, making it suitable for high-frequency creation, batch concept exploration, and design iteration. On this platform, generation, reference image editing, and result delivery use the same image interface and can be integrated into daily content production workflows.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
The platform supports 1–3 reference images and outputs 1–6 images per request
Native output dimensions
Total pixel area from 512×512 to 2048×2048; aspect ratios from 1:8 to 8:1
Prompt input
Supports Chinese and English; platform length limit of 18,000 characters
Enhancement and reasoning
direct supports generation and editing; agent supports text-to-image only; reasoning mode requires prompt enhancement to be enabled
Image controls
Negative prompts, random seed, watermark toggle
Delivery methods
Synchronous or asynchronous tasks, completion callbacks, CDN image URLs
Dimensions are native model specifications; reference image submission, dimension syntax, and result delivery follow this platform's interface rules.
Core Capabilities
From Creative Descriptions to Visual Concepts
Describe the subject, composition, materials, lighting, and style in Chinese or English to generate images directly. Ideal for creative workflows that first explore visual directions and then gradually refine requirements. Prompt enhancement can expand simple descriptions, and negative prompts can specify content you do not want to appear, keeping iterations focused on clear visual goals.
Use Reference Images to Clarify What to Edit
Provide reference images and write editing instructions to modify the background, colors, style, or visual elements. For multi-image tasks, explain the purpose of each image in order, such as subject reference, composition reference, or color reference, while listing features that need to be preserved to reduce confusion between different reference inputs.
Batch Exploration and Task-Based Delivery
Generate multiple candidate images in a single request to compare compositions and visual directions. Interactive creation can use synchronous calls, while background production can use asynchronous tasks and completion callbacks. Successful results provide CDN image URLs, allowing business systems to continue with previewing, downloading, and distribution without tying the generation process to the display page.
Use Cases
Exploring Marketing Visual Directions
Enter campaign themes, product features, target audiences, color palettes, and whitespace requirements to generate promotional image or e-commerce visual candidates. First compare different options in batches, then refine the background and lighting for the selected direction. The deliverables are suitable as initial design drafts; final copy and layouts can still be organized in design tools.
Scene Redesigns for Existing Assets
Submit product photos or character assets, specifying the background, color tone, and style to be changed, as well as subject features that must be retained. When multiple references are needed, provide separate subject and style images and clearly define their respective roles. The output can be used for scene concept reviews, with review focusing on whether subject details are retained and whether changes follow the instructions.
Ongoing Image Creation in Content Systems
Organize section themes, visual guidelines, and size requirements into prompts to continuously generate image candidates. The backend retrieves results through asynchronous tasks, then passes image URLs to the content management system. This is suitable for production workflows with human image selection, allowing generation, review, and publishing to be separated and preventing unreviewed images from entering official content directly.
How to choose this model
The Standard version is suitable for broad exploration first
When tasks require frequent prompt adjustments, comparing multiple compositions, or continuously producing candidate images, prioritize qwen-image-3.0. It is positioned to balance quality and speed, making it suitable for screening directions before refining. Compared with qwen-image-3.0-pro, the key consideration is iteration efficiency rather than assuming every image can be delivered directly as a high-precision final output.
Consider Pro for complex final images
If the goal is complex layouts, commercial posters, or final visuals that place greater emphasis on detail, consider qwen-image-3.0-pro from the same series. Both support generation and editing; the difference is not whether reference images can be used. It is recommended to compare final image results using the same materials and acceptance criteria, then choose according to the actual task rather than treating the version name as a quality guarantee in all scenarios.
Get started
Prepare copy and reference roles
Organize the title, body text, and layout; when editing is needed, prepare 1–3 public web images and separately specify the subject, background, and style references.
Clearly choose Standard or Pro
Send model=qwen-image-3.0 and prompt to /qwen-image/images; start with size=1024*1024 and n=1. Use image_urls when editing, and do not write dimensions with x as the separator.
Retrieve generated images and continue revising
Read image results synchronously, or use async=true to save task_id and then query /qwen-image/tasks; verify the text and subject, then use the selected image for the next round of editing.
Trial suggestion: two layouts for an event image
Input and goal
Create a square image for a community reading event: an open book by a window, warm morning light, space in the upper right for the title “Read Together This Weekend,” and simple visual elements.
Acceptance and next steps
Check the title, book pages, and whitespace; compare the two layouts, select one, then edit using the original image, avoiding changing both the subject and style in the same round.
Usage Boundaries
Multi-reference image editing requires clearly specifying the role of each image. If requirements for the subject, style, and composition conflict with one another, the result may blend attributes you do not wish to retain. When involving a person's appearance, product structure, or brand identity, check each item individually; do not treat an instruction to “preserve the subject” as a guarantee of pixel-level consistency.
agent prompt enhancement applies only to text-to-image generation; reference image editing should use direct. Thinking mode requires prompt enhancement to be enabled at the same time and increases generation time; therefore, complex instructions and rapid experimentation should use different control combinations, and it is not advisable to assume that more enhancement options are always better.
Output dimensions are constrained by both total pixel area and aspect ratio, so not every width and height can be generated. Images containing text, annotations, or infographic elements should be checked separately; image results are also not editable layout files, so post-production is still needed when precise font sizes and layouts are required.
Frequently Asked Questions
Can qwen-image-3.0 directly modify existing images?
Yes. Provide 1–3 publicly accessible image URLs, and specify the modifications and preservation requirements in the prompt to perform reference image editing. Without reference images, it is used for text-to-image generation. For multiple reference images, explain their purposes in submission order to avoid listing images without explaining their respective roles.
How should dimensions be specified when calling it?
The size on this platform uses the width*height format, such as 1024*1024, rather than 1024x1024. When selecting dimensions, you must also meet the native total pixel area and aspect ratio limits; these ranges do not mean that every side must be between 512 and 2048 pixels.
How should prompt enhancement and thinking mode be combined?
direct can be used for text-to-image generation and reference image editing, while agent is only for text-to-image generation. When enable_thinking is enabled, prompt_extend=true is required. Simple concept exploration can reduce enhancement steps, while complex images can try enhancement and thinking, but the additional generation wait time should also be considered.
How should I choose between the Standard version and Pro?
The Standard version is suitable for batch creation, rapid iteration, and everyday visual exploration; Pro is better suited for complex layouts and high-precision final-image requirements. It is recommended to compare detail performance and the number of revisions for the same task, rather than looking only at a single showcase image. They are different models, and Pro's name should not be treated as an alias for the Standard version.
How do I retrieve images after asynchronous generation?
After setting async=true, save the returned task_id and query the final task status through /qwen-image/tasks, or use callback_url to receive a completion notification. After success, read image_url from the result; images will be transferred to this platform's CDN and can be used for subsequent display, download, and distribution.