Event date · · Tencent

Tencent open-sources WeMM-Embedding-4B multimodal embedding model

Tencent 腾讯Chinese AIOpen weights
FACT STATEMENT

Tencent released WeMM-Embedding-4B, a universal multimodal embedding model built on Qwen3.5, on Hugging Face under Apache License 2.0. It accepts text, images, videos, visual documents, and interleaved inputs, returning 2,560-dimensional L2-normalized embeddings. The model achieves 79.2 average on MMEB-v2 and 58.2 on MMEB-v3.

China context

Original name
腾讯
Outside China
Open weights · huggingface.co
Claims
Company-reported; not yet independently evaluated
For builders
Builders outside China can integrate a permissively licensed multimodal embedding model for retrieval, search, and agent memory without API costs or data leaving their infrastructure.
For investors
Tencent's release of competitive open-weights embedding models signals continued investment in open-source AI, potentially commoditizing embedding APIs and shifting value to application-layer differentiation.
What happened

Tencent published WeMM-Embedding-4B on Hugging Face, a 4-billion-parameter multimodal embedding model built on Qwen3.5. It supports text, images, videos, visual documents, and interleaved multimodal inputs, producing 2,560-dimensional L2-normalized embeddings. The model is licensed under Apache License 2.0 and can be used with Transformers, Sentence Transformers, vLLM, and SGLang. Evaluation results on MMEB-v2 (78 datasets) show an average score of 79.2, outperforming smaller WeMM-Embedding 2B (77.9) and Qwen3-VL-Embedding 2B (73.2). On MMEB-v3 (190 tasks), it achieves 58.2 overall, with 47.9 on text tasks and 49.0 on agent tasks.

Technical significance

WeMM-Embedding-4B is a 4B-parameter multimodal embedding model based on Qwen3.5, supporting text, image, video, visual document, and interleaved inputs. It outputs 2,560-dimensional L2-normalized embeddings and supports Matryoshka embeddings via truncate_dim. The model can be served with vLLM 0.27.0 and SGLang 0.5.9. On MMEB-v2, it achieves 79.2 average (image 80.8, video 72.1, visual document 82.0), and on MMEB-v3, 58.2 overall (text 47.9, agent 49.0, MCMR 41.9).

Industry impact

Developers outside China can now use a competitive open-weights multimodal embedding model from Tencent under Apache 2.0, reducing reliance on proprietary APIs for retrieval and multimodal search. This directly pressures closed-source embedding providers and expands the open-weights ecosystem for multimodal RAG and agent applications.

Decision value

For enterprises building multimodal search, recommendation, or retrieval systems, WeMM-Embedding-4B offers a permissively licensed, state-of-the-art embedding model that can be self-hosted, avoiding per-query API costs and data privacy concerns. Its strong performance on visual documents and agent tasks makes it suitable for document understanding and agentic workflows.

What to watch

Next signals to check: whether Tencent releases larger WeMM-Embedding variants (e.g., 9B) with open weights, whether the model appears in third-party benchmarks or leaderboards, and whether community fine-tunes or integrations emerge on Hugging Face.

Latest in Chinese AI

  1. MiniMaxMiniMax open-sources MiniMax-Code-MiniApps repository for community-built plugins
  2. Manus AIManus AI Launches Video Editor in Manus 2.0
  3. DeepSeekDeepSeek open-sources dsh-libreoffice-kit 0.1.0 for font-friendly Office conversion and rendering in Node.js
  4. DeepSeekDeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
  5. Beijing Academy of Artificial IntelligenceBAAI released AREX-2, a 27B self-improving agent model, on Hugging Face

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.