Alibaba PAI released a unified ControlNet for Qwen-Image 2.1 covering 8 conditions and inpainting
Alibaba PAI released Qwen-Image-2.1-Fun-Controlnet-Union on Hugging Face, a ControlNet-Union branch for Qwen-Image 2.1 that handles 8 control conditions and inpainting in one checkpoint. The checkpoint is about 7.0 GB and loads on top of the base Qwen-Image 2.1 transformer.
China context
- Original name
- 阿里巴巴
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers can download the checkpoint from Hugging Face and integrate it with Qwen-Image 2.1 for controlled generation and inpainting without per-condition weights.
- For investors
- The release extends Alibaba's open-weights image generation ecosystem, but the 'other' license may limit commercial adoption until clarified.
Alibaba PAI released Qwen-Image-2.1-Fun-Controlnet-Union on Hugging Face. It is a ControlNet-Union branch for Qwen-Image 2.1, a flow-matching text-to-image DiT. A single checkpoint drives 8 structural control conditions (Canny, Depth, Grayscale, HED, Lineart, MLSD, Pose, Scribble) and image inpainting, without per-condition weights. The checkpoint holds only the control branch (control img in plus 16 control blocks, about 7.0 GB) and is loaded with strict=False on top of the base Qwen-Image 2.1 transformer. The control branch attaches a skip to every 2nd of the 32 transformer blocks (16 injection points), with zero-gated before proj / after proj projections. Control and inpainting share one branch: control input is widened to control in dim = 129 (control latents 64, mask 1, masked-image latents 64). For pure control the mask/masked-image channels are zero-padded; for inpainting the same branch re-draws the masked region from the prompt. The two can be combined. Example scripts run with guidance scale = 1.0 (CFG-distilled fast sampling). control context scale scales every control skip before it is added to the main branch: 1.0 is strongest, 0.0 switches control off. Qwen-Image 2.1 encodes prompt and condition image with a Qwen3-VL text encoder + processor, and its VAE decodes to RGBA, so previews are saved as PNG. Samples in the model card are generated with num inference steps = 40, control context scale = 1.0, seed 43.
The ControlNet-Union branch uses dense control injection: a skip is attached to every 2nd of the 32 transformer blocks (16 injection points), and each control skip is added back to the main branch through zero-gated before proj / after proj projections. The control input is widened to control in dim = 129 to accommodate control latents (64), mask (1), and masked-image latents (64), enabling control and inpainting to share one branch. The model runs with guidance scale = 1.0, indicating CFG-distilled fast sampling.
Developers outside China can now use a single open-weights checkpoint for 8 control conditions and inpainting with Qwen-Image 2.1, reducing the need to manage multiple per-condition models. This lowers integration cost for image generation pipelines that require structural control.
The release provides a unified control mechanism for Qwen-Image 2.1, potentially simplifying workflows for developers who need multiple control types. However, the license is listed as 'other', so commercial use may require clarification.
A specific signal to check is whether the model card or repository adds a license file or clarifies the 'other' license, and whether independent benchmarks compare its control adherence against per-condition ControlNets.