InternLM released InternLumina-U2, a unified multimodal model for text, image, video, and 3D
InternLM released InternLumina-U2, a 16B-parameter MoE with 1B active parameters, on Hugging Face under Apache 2.0. It handles text QA, image/video/3D understanding, image generation and editing. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs.
China context
- Original name
- 上海人工智能实验室
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Builders outside China can download the Apache 2.0 weights from Hugging Face and run inference on NVIDIA GPUs using the provided code, with the option to switch to Huawei Ascend NPUs if needed.
- For investors
- The release of a unified multimodal model with dual-hardware support signals that Shanghai AI Lab is targeting both international and domestic deployment markets, which may affect competitive dynamics for Western multimodal model providers.
InternLM released InternLumina-U2, a unified multimodal model that brings language, image, video and 3D into a single framework. It is a 16B-parameter MoE with 1B active parameters (16B-A1B), pairing an efficient sparse backbone with an 8-codebook fully-discrete visual representation built on AToken. The model covers text QA, text-to-image generation, image understanding, image editing, video understanding and 3D understanding with one model. Checkpoints are hosted on Hugging Face under Apache 2.0, with separate folders for Huawei Ascend NPUs and NVIDIA GPUs. Inference code and examples are available on GitHub.
InternLumina-U2 uses a sparse Mixture-of-Experts architecture with 16B total parameters and 1B active parameters, paired with an 8-codebook fully-discrete visual representation built on AToken. This design aims to unify multiple modalities—text, image, video, and 3D—within a single framework. The model is released with checkpoints for both Huawei Ascend NPUs and NVIDIA GPUs, indicating compatibility with domestic Chinese hardware.
Developers outside China can now access a unified multimodal model that runs on both NVIDIA and Huawei Ascend hardware, reducing the cost of switching between separate models for understanding and generation tasks. The Apache 2.0 license removes legal friction for commercial use, while the dual-hardware checkpoints give teams flexibility in deployment environments.
The Apache 2.0 license permits commercial use, and the availability of checkpoints for both NVIDIA and Huawei Ascend hardware may lower deployment costs for enterprises with existing infrastructure on either stack. The unified model could reduce the need to maintain separate models for understanding and generation tasks.
The technical report is marked as 'Coming Soon', and full benchmark comparison tables will appear there. Observers can check the GitHub repository for the current preliminary benchmark table and watch for the report's release to assess performance relative to other multimodal models.