Tencent open-sourced WeMM-Embedding-9B, a multimodal embedding model
Tencent released WeMM-Embedding-9B, a universal multimodal embedding model built on Qwen3.5, on Hugging Face. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, returning 4,096-dimensional L2-normalized embeddings. The model is licensed under Apache License 2.0.
China context
- Original name
- 腾讯
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Builders outside China can integrate WeMM-Embedding-9B via Hugging Face Transformers or Sentence Transformers, and serve it with vLLM or SGLang, using Apache 2.0 licensed weights.
- For investors
- Tencent's release of a competitive open-weight multimodal embedding model under Apache 2.0 may pressure commercial embedding API providers and signal continued investment in open-source AI by Chinese tech giants.
Tencent released WeMM-Embedding-9B, a universal multimodal embedding model built on Qwen3.5, on Hugging Face. The model accepts text, images, videos, visual documents, and interleaved multimodal inputs, returning 4,096-dimensional L2-normalized embeddings. Audio input is not supported. The model is licensed under Apache License 2.0. Evaluation results on MMEB-v2 and MMEB-v3 benchmarks are provided in the model card.
WeMM-Embedding-9B is a 9B parameter multimodal embedding model built on Qwen3.5. It supports text, image, video, visual document, and interleaved multimodal inputs, producing 4,096-dimensional L2-normalized embeddings. The model supports Matryoshka embeddings via truncate_dim with renormalization. It can be served with vLLM 0.27.0 or SGLang 0.5.9. On MMEB-v2 (78 datasets), it achieves an average score of 80.6, with image Hit@1 of 81.9, video Hit@1 of 74.3, and visual document NDCG@5 of 83.3. On MMEB-v3 (190 tasks), it achieves V3-All score of 59.5, text NDCG@5 of 48.8, agent Hit@1 of 51.0, MCMR Hit@1 of 49.3, and audio Hit@1 of 0.0 (audio unsupported).
Developers outside China can now use a state-of-the-art open-weight multimodal embedding model from Tencent under Apache 2.0, reducing the cost and complexity of building retrieval systems for mixed text, image, and video content compared to closed-source alternatives.
WeMM-Embedding-9B provides a permissively licensed, high-performance embedding model for multimodal retrieval, enabling businesses to build search and recommendation systems over diverse content types without vendor lock-in.
Observable next signals include adoption of WeMM-Embedding-9B in open-source retrieval pipelines, community benchmarks comparing it to Qwen3-VL-Embedding and other models, and potential releases of larger or audio-capable variants by Tencent.