Ant Group's inclusionAI releases Ming-Image open-source visual design models
Ant Group's inclusionAI released Ming-Image-0.1-Design and Ming-Image-0.1-Design-Layer, two 6B-parameter open-source models for visual design generation and editable layer decomposition. The models are available on GitHub, Hugging Face, and ModelScope.
China context
- Original name
- 蚂蚁集团
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers can integrate the models locally using the provided CLI and Python downloader, with checkpoints available on Hugging Face and ModelScope. The layer decomposition model outputs transparent RGBA layers, useful for design-to-code workflows.
- For investors
- Ant Group's open-source release of design-focused models signals investment in creative AI tools, potentially competing with proprietary design software. The 80 GiB GPU requirement may limit adoption to well-resourced teams.
inclusionAI, affiliated with Ant Group, published the Ming-Image repository containing two 6B-parameter models: Ming-Image-0.1-Design for generating complete visual designs (UI, infographics, posters, text-rich compositions) and Ming-Image-0.1-Design-Layer for decomposing flattened designs into editable transparent layers. The repository includes inference code, prompt rewriting guidance, and integration with agent workflows such as Ling UI Design and Image to Editable PPT. Requirements include Python 3.10+, CUDA-capable PyTorch, and at least one GPU with 80 GiB memory for default deployment.
The models use a diffusion transformer with PyTorch SDPA internally and support eager or FlashAttention 2 attention implementations. Text-to-image defaults to 2048x2048 resolution with 12 steps and CFG 1.0; layer decomposition uses 12 steps and CFG 2.0 with a working resolution of 512 or 1024. The layer model returns requested layers plus a composite image, and the CLI skips the composite. Prompt rewriting is handled by an instruction-following VLM (Ling-3.0-flash-VL or qwen3.8-27B) as a preprocessing step.
Designers and developers outside China can now use these open-weight models to generate and decompose visual designs locally, reducing reliance on proprietary design tools. The release of layer decomposition capability directly addresses editable output, a gap in many text-to-image models.
The open-source release enables enterprises to integrate visual design generation and layer decomposition into their pipelines without per-image API costs. The 80 GiB GPU requirement may limit deployment to high-end hardware, but the availability of Hugging Face and ModelScope checkpoints lowers the barrier for experimentation.
Next signals to watch: adoption of the models in agent workflows (Ling UI Design, Image to Editable PPT), community contributions to the repository, and whether inclusionAI releases larger or specialized variants. Verification needed on whether the models are accessible via API outside China or only as open weights.