InternLM open-sources InstantFusion, a shared latent interface for composing heterogeneous image generators
InternLM (Shanghai AI Laboratory) released the official implementation of InstantFusion, a shared latent interface for heterogeneous image generators, on GitHub. The repository includes training and inference code for latent alignment, cross-model on-policy distillation, and model handoff. Trained checkpoints are available on Hugging Face.
China context
- Original name
- 上海人工智能实验室
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can access the code and checkpoints on GitHub and Hugging Face, but must provide their own base model weights (Qwen-Image, SD3) and training data (VisPrompt5M), which may require separate licensing.
- For investors
- The release signals Shanghai AI Laboratory's continued investment in open-source multimodal infrastructure, potentially increasing competition in the image generation tooling space.
InternLM (Shanghai AI Laboratory) open-sourced InstantFusion, a method that maps latents of different diffusion models into a shared space, enabling mid-denoising handoff between models. The repository provides code for latent alignment encoder training, cross-model on-policy distillation, and LAE-based inference. It supports acceleration, multi-preference composition, and cross-model distillation. Checkpoints are hosted on Hugging Face, and the paper is forthcoming.
InstantFusion uses sigma-conditioned encoders and decoders to map latents from different diffusion models into a shared space, trained with reconstruction, cross-reconstruction, latent alignment, and denoising-velocity alignment losses. Cross-model OPD trains an SD3 LoRA with Qwen-Image as a frozen teacher, supervising zero-based steps 1,2,4,8 by default. The repository includes scripts for LAE training, OPD training, and LAE-based inference, with configuration via JSON files.
Developers working with multiple diffusion models can now compose them without retraining from scratch, reducing the cost of combining specialized models for different preferences or speed/quality trade-offs. This directly affects teams using Qwen-Image, FLUX.1, or SD3, as they can hand off denoising trajectories between models. The open-source release may pressure other labs to provide similar interoperability layers.
For enterprises using multiple image generation models, InstantFusion could reduce infrastructure costs by enabling model handoff for acceleration or preference composition without maintaining separate pipelines. However, the repository does not include model weights or training data, so users must supply their own Qwen-Image and SD3 checkpoints and VisPrompt5M dataset, which may limit immediate adoption.
The paper is marked as 'coming soon'; its release will clarify the method's theoretical grounding and benchmark results. Adoption can be tracked via GitHub stars, forks, and issues on the InstantFusion repository, as well as community projects building on the shared latent space. The availability of checkpoints on Hugging Face may lead to third-party fine-tunes or integrations.