Tencent open-sources WeMM-Embedding-4B multimodal embedding model
Tencent released WeMM-Embedding-4B, a universal multimodal embedding model built on Qwen3.5, on Hugging Face under Apache License 2.0. It accepts text, images, videos, visual documents, and interleaved inputs, returning 2,560-dimensional L2-normalized embeddings. The model achieves 79.2 average on MMEB-v2 and 58.2 on MMEB-v3.
China context
- Original name
- 腾讯
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Builders outside China can integrate a permissively licensed multimodal embedding model for retrieval, search, and agent memory without API costs or data leaving their infrastructure.
- For investors
- Tencent's release of competitive open-weights embedding models signals continued investment in open-source AI, potentially commoditizing embedding APIs and shifting value to application-layer differentiation.
Tencent published WeMM-Embedding-4B on Hugging Face, a 4-billion-parameter multimodal embedding model built on Qwen3.5. It supports text, images, videos, visual documents, and interleaved multimodal inputs, producing 2,560-dimensional L2-normalized embeddings. The model is licensed under Apache License 2.0 and can be used with Transformers, Sentence Transformers, vLLM, and SGLang. Evaluation results on MMEB-v2 (78 datasets) show an average score of 79.2, outperforming smaller WeMM-Embedding 2B (77.9) and Qwen3-VL-Embedding 2B (73.2). On MMEB-v3 (190 tasks), it achieves 58.2 overall, with 47.9 on text tasks and 49.0 on agent tasks.
WeMM-Embedding-4B is a 4B-parameter multimodal embedding model based on Qwen3.5, supporting text, image, video, visual document, and interleaved inputs. It outputs 2,560-dimensional L2-normalized embeddings and supports Matryoshka embeddings via truncate_dim. The model can be served with vLLM 0.27.0 and SGLang 0.5.9. On MMEB-v2, it achieves 79.2 average (image 80.8, video 72.1, visual document 82.0), and on MMEB-v3, 58.2 overall (text 47.9, agent 49.0, MCMR 41.9).
Developers outside China can now use a competitive open-weights multimodal embedding model from Tencent under Apache 2.0, reducing reliance on proprietary APIs for retrieval and multimodal search. This directly pressures closed-source embedding providers and expands the open-weights ecosystem for multimodal RAG and agent applications.
For enterprises building multimodal search, recommendation, or retrieval systems, WeMM-Embedding-4B offers a permissively licensed, state-of-the-art embedding model that can be self-hosted, avoiding per-query API costs and data privacy concerns. Its strong performance on visual documents and agent tasks makes it suitable for document understanding and agentic workflows.
Next signals to check: whether Tencent releases larger WeMM-Embedding variants (e.g., 9B) with open weights, whether the model appears in third-party benchmarks or leaderboards, and whether community fine-tunes or integrations emerge on Hugging Face.