Event date · · Ant Group

inclusionAI released Qwen3.8-27B-singprobe, a streaming guardrail probe for Qwen3.8-27B

Ant Group 蚂蚁集团Chinese AIOpen weights
FACT STATEMENT

inclusionAI released Qwen3.8-27B-singprobe, an Apache-2.0 licensed streaming guardrail probe built on Qwen/Qwen3.8-27B, on Hugging Face. It adds less than 0.5% decode-time overhead and scores query intent, response unsafety, and hallucination risk at every token.

China context

Original name
inclusionAI
Outside China
Open weights · huggingface.co
Claims
Company-reported; not yet independently evaluated
For builders
Developers outside China can download the model weights from Hugging Face and integrate the probe via SGLang or vLLM branches, enabling streaming safety checks without a separate model.
For investors
The release of a low-overhead safety probe by an Ant Group-affiliated entity may indicate a strategic focus on enterprise AI safety, potentially affecting the competitive landscape for AI safety startups.
What happened

inclusionAI released Qwen3.8-27B-singprobe on Hugging Face under Apache-2.0. The model is an intrinsic streaming guardrail that reuses the base model's hidden states to score query intent, response unsafety, and hallucination risk at every token, adding less than 0.5% decode-time overhead. It uses 10.1M probe parameters tapped at layers 20, 41, and 62, and outputs 8 intents plus unsafe and hallucination scores. Evaluation results show F1 of 0.8683 on query intent classification, 0.8676 on response safety classification, R-AUC/T-AUC of 0.9881/0.9305 on streaming safety, and AUC of 0.8035 on hallucination detection. The model is supported through SGLang and vLLM integration branches.

Technical significance

SingProbe is a lightweight probe (10.1M parameters) that taps hidden states from layers 20, 41, and 62 of Qwen3.8-27B to produce per-token scores for 8 intents, unsafety, and hallucination. It achieves a benign-response false-positive rate of 0.03% and less than 0.5% decode overhead. The approach avoids running a separate safety model, reducing latency and compute costs for streaming guardrails.

Industry impact

Developers using Qwen3.8-27B can add streaming safety and hallucination detection with minimal overhead, reducing the need for separate guardrail models and lowering inference costs. This may pressure providers of standalone guardrail models to offer tighter integration or lower prices.

Decision value

The model offers a cost-effective safety solution for enterprises deploying Qwen3.8-27B, potentially reducing compliance overhead and latency. Its Apache-2.0 license allows commercial use, which may accelerate adoption in production systems.

What to watch

Adoption can be tracked by monitoring downloads and community feedback on the Hugging Face model page, as well as pull requests or issues in the SGLang and vLLM integration branches. Independent benchmarks comparing SingProbe against reference baselines on the same datasets would validate the reported metrics.

Latest in Chinese AI

  1. MiniMaxMiniMax open-sources MiniMax-Code-MiniApps repository for community-built plugins
  2. Manus AIManus AI Launches Video Editor in Manus 2.0
  3. DeepSeekDeepSeek open-sources dsh-libreoffice-kit 0.1.0 for font-friendly Office conversion and rendering in Node.js
  4. DeepSeekDeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
  5. Beijing Academy of Artificial IntelligenceBAAI released AREX-2, a 27B self-improving agent model, on Hugging Face

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.