Event date · · Ant Group

inclusionAI open-sourced Qwen3.5-27B-singprobe, a streaming safety and hallucination guardrail for Qwen3.5-27B

Ant Group 蚂蚁集团Chinese AIOpen weights
FACT STATEMENT

inclusionAI released Qwen3.5-27B-singprobe, an Apache-2.0 licensed streaming guardrail probe built on Qwen/Qwen3.5-27B, on Hugging Face. It adds less than 0.5% decode-time overhead and scores query intent, response unsafety, and hallucination risk at every token.

China context

Original name
inclusionAI
Outside China
Open weights · huggingface.co
Claims
Company-reported; not yet independently evaluated
For builders
Developers outside China can download the Apache-2.0 weights from Hugging Face and integrate the probe via SGLang or vLLM branches to add streaming safety and hallucination detection to Qwen3.5-27B deployments.
For investors
Ant Group's inclusionAI is releasing open-weights safety tooling, which may increase adoption of Qwen models in regulated industries and create demand for integration and evaluation services.
What happened

inclusionAI released Qwen3.5-27B-singprobe on Hugging Face under Apache-2.0. The model is an intrinsic streaming guardrail that reuses the base model's hidden states to score query intent, response unsafety, and hallucination risk at every token, adding less than 0.5% decode-time overhead. It uses 10.1M probe parameters tapped at layers 20, 41, and 62, and outputs 8 intents plus unsafe and hallucination scores. Evaluation results show F1 of 0.8741 on query intent classification, 0.8729 on response safety classification, R-AUC/T-AUC of 0.9879/0.9289 on streaming safety, and AUC of 0.7967 on hallucination detection. The model is supported through SGLang and vLLM integration branches.

Technical significance

SingProbe is a lightweight intrinsic guardrail that reuses the base model's hidden states during generation, avoiding a separate safety model. It taps layers 20, 41, and 62 with 10.1M probe parameters and outputs per-token scores for 8 intents, unsafe, and hallucination. Reported decode overhead is under 0.5%, and benign-response false-positive rate is 0.03% across 5 datasets. The model is available on Hugging Face with Apache-2.0 license, and training code is at inclusionAI/SingProbe.

Industry impact

Developers using Qwen3.5-27B can add streaming safety and hallucination detection with minimal latency overhead, reducing the need for separate guardrail models and potentially lowering inference costs. This release positions Ant Group's inclusionAI as a provider of open-weights safety tooling for the Qwen ecosystem.

Decision value

For enterprises deploying Qwen3.5-27B, this probe offers a low-overhead way to add content safety and hallucination detection without a separate model, potentially reducing infrastructure costs and simplifying compliance. The Apache-2.0 license allows commercial use.

What to watch

Next signals to check: adoption of the SGLang/vLLM integration branches, independent benchmarks of the probe's false-positive rate and overhead, and whether inclusionAI releases probes for other base models or expands the intent taxonomy.

Latest in Chinese AI

  1. MiniMaxMiniMax open-sources MiniMax-Code-MiniApps repository for community-built plugins
  2. Manus AIManus AI Launches Video Editor in Manus 2.0
  3. DeepSeekDeepSeek open-sources dsh-libreoffice-kit 0.1.0 for font-friendly Office conversion and rendering in Node.js
  4. DeepSeekDeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
  5. Beijing Academy of Artificial IntelligenceBAAI released AREX-2, a 27B self-improving agent model, on Hugging Face

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.