Event date · · Ant Group

inclusionAI open-sourced Qwen3.5-9B-singprobe, a streaming guardrail probe for Qwen3.5-9B

Ant Group 蚂蚁集团Chinese AIOpen weights
FACT STATEMENT

inclusionAI released Qwen3.5-9B-singprobe, an Apache-2.0 licensed streaming guardrail probe built on Qwen/Qwen3.5-9B, on Hugging Face. It adds less than 0.5% decode-time overhead and scores query intent, response unsafety, and hallucination risk at every token.

China context

Original name
inclusionAI
Outside China
Open weights · huggingface.co
Claims
Company-reported; not yet independently evaluated
For builders
Developers outside China can download the model from Hugging Face and integrate it via SGLang or vLLM to add streaming safety and hallucination detection to Qwen3.5-9B deployments.
For investors
The release of an open-weights guardrail probe by an Ant Group-affiliated entity may indicate a strategic move to strengthen the Qwen ecosystem and compete with dedicated safety model providers.
What happened

inclusionAI released Qwen3.5-9B-singprobe on Hugging Face under Apache-2.0. The model is an intrinsic streaming guardrail built on Qwen/Qwen3.5-9B that reuses the base model's hidden states to score query intent, response unsafety, and hallucination risk at every token, adding less than 0.5% decode-time overhead. It has 8.13M probe parameters, taps layers [9, 19, 30], and outputs 8 intents plus unsafe and hallucination scores. Evaluation results show F1 of 0.8719 on query intent classification, 0.8695 on response safety classification, R-AUC/T-AUC of 0.9874/0.9344 on streaming safety, and AUC of 0.7954 on hallucination detection. The model is supported through SGLang and vLLM integration branches.

Technical significance

SingProbe is a lightweight probe (8.13M parameters) that taps hidden states from layers 9, 19, and 30 of Qwen3.5-9B to produce per-token scores for 8 intents, unsafety, and hallucination risk. It adds less than 0.5% decode-time overhead and achieves a benign-response false-positive rate of 0.04% across 5 datasets. The model card reports benchmark results that are close to or slightly below dedicated guard models, but with much lower overhead.

Industry impact

Developers using Qwen3.5-9B can add streaming safety and hallucination detection without deploying a separate guard model, reducing infrastructure cost and latency. This may pressure dedicated guardrail model providers, as the probe is open-weights and Apache-2.0 licensed.

Decision value

For enterprises deploying Qwen3.5-9B, SingProbe offers a low-overhead way to add safety and hallucination monitoring, potentially reducing the need for separate guard models and lowering operational costs.

What to watch

Adoption can be tracked by monitoring downloads and community usage of the Hugging Face model, as well as activity in the SGLang and vLLM integration branches. Independent evaluation of the reported benchmarks would verify the performance claims.

Latest in Chinese AI

  1. MiniMaxMiniMax open-sources MiniMax-Code-MiniApps repository for community-built plugins
  2. Manus AIManus AI Launches Video Editor in Manus 2.0
  3. DeepSeekDeepSeek open-sources dsh-libreoffice-kit 0.1.0 for font-friendly Office conversion and rendering in Node.js
  4. DeepSeekDeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
  5. Beijing Academy of Artificial IntelligenceBAAI released AREX-2, a 27B self-improving agent model, on Hugging Face

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.