Tencent open-sourced WeMM-Embedding-2B, a multimodal embedding model
Tencent released WeMM-Embedding-2B, a universal multimodal embedding model built on Qwen3.5, on Hugging Face under Apache License 2.0. It accepts text, images, videos, visual documents, and interleaved inputs, returning 2,048-dimensional embeddings. On MMEB-v2, it achieves an average score of 77.9 across 78 datasets.
China context
- Original name
- 腾讯
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can integrate WeMM-Embedding-2B via Hugging Face Transformers or Sentence Transformers, with support for vLLM and SGLang serving. The model's Matryoshka embeddings allow dimension reduction to 256 with minimal performance loss, enabling efficient large-scale retrieval.
- For investors
- Tencent's release of a competitive open-weight multimodal embedding model signals continued investment in open-source AI, which may pressure other model providers to release or improve their own embedding models. The model's strong benchmark results could drive adoption in enterprise search and RAG applications.
WeMM-Embedding-2B is a 2-billion-parameter multimodal embedding model from Tencent, built on Qwen3.5. It supports text, images, videos, visual documents, and interleaved multimodal inputs, producing 2,048-dimensional L2-normalized embeddings. Audio is not supported. The model is available on Hugging Face under the Apache License 2.0, with support for Transformers, Sentence Transformers, vLLM, and SGLang. It supports Matryoshka embeddings, with 256-dimensional embeddings retaining 98.7% of full-dimensional image and video performance on MMEB-v2. On MMEB-v2, it outperforms other 2B models with an average score of 77.9, and on MMEB-v3 it scores 56.0 on all 190 tasks.
WeMM-Embedding-2B uses a 2,048-dimensional embedding space and supports Matryoshka dimensionality reduction, allowing efficient storage and retrieval with minimal performance loss. It is built on Qwen3.5 and accepts interleaved multimodal inputs via chat message format. The model does not support audio input, which is reflected in zero scores on audio tasks in MMEB-v3.
Developers outside China can now use a state-of-the-art open-weight multimodal embedding model from Tencent, reducing the cost and complexity of building retrieval systems for mixed text, image, and video content. This directly competes with other open multimodal embedding models like Qwen3-VL-Embedding and GME, potentially shifting adoption toward Tencent's offering.
WeMM-Embedding-2B provides a free, open-weight alternative for multimodal retrieval, potentially reducing infrastructure costs for businesses that need to search across text, images, and videos. Its strong benchmark performance may make it a default choice for RAG and multimodal search applications.
The next observable signal is whether Tencent releases the larger 4B and 9B variants of WeMM-Embedding on Hugging Face, as they are mentioned in the model card but not yet available. Another signal is community adoption, measurable by downloads, GitHub stars, and integration into popular retrieval frameworks.