Qwen open-sourced Qwen-Image-2.1-Turbo, an 8-step accelerated image model
Qwen released Qwen-Image-2.1-Turbo, an accelerated checkpoint of Qwen-Image-2.1 for text-to-image generation and image editing with 8 denoising steps. It uses the same 7B visual generation architecture and loads directly with QwenImage21Pipeline in Diffusers. The checkpoint includes its recommended sampling schedule and is available on Hugging Face.
China context
- Original name
- 通义千问图像2.1 Turbo
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can integrate the Turbo checkpoint via Diffusers for faster text-to-image and image editing, but must review the Qwen Research License Agreement for commercial use restrictions.
- For investors
- The release of an accelerated checkpoint signals Qwen's focus on inference efficiency, which may affect competitive positioning in the open-weights image model market.
Qwen-Image-2.1-Turbo is an accelerated checkpoint of Qwen-Image-2.1 for text-to-image generation and image editing with 8 denoising steps. It uses the same 7B visual generation architecture and loads directly with QwenImage21Pipeline in Diffusers. The checkpoint includes its recommended sampling schedule, so it is ready to use without manually configuring the scheduler. Generation uses CFG=1 by default, and prefix KV caching reuses the text and reference-image context across denoising steps. The model is licensed under the Qwen Research License Agreement.
The checkpoint requires Diffusers with support for pipeline-configured sampling sigmas, added in PR #14950. The recommended 8-step sampling schedule is saved with the checkpoint and loaded automatically; setting num inference steps alone does not override it. An explicit call-time sigmas argument overrides the saved schedule, but other schedules have not been evaluated for this checkpoint.
Developers outside China can now use a faster Qwen image model with 8-step sampling, reducing inference cost and latency for text-to-image and image editing tasks. This puts pressure on other open-weights image models to offer similar speed optimizations.
The Turbo checkpoint offers faster inference with 8 denoising steps, potentially lowering compute costs for developers integrating text-to-image and image editing capabilities.
Observable next signals include adoption of the Turbo checkpoint in downstream applications, community benchmarks comparing quality and speed against other accelerated image models, and any updates to the Qwen Research License Agreement that affect commercial use.