Event date · · Ant Group

inclusionAI open-sourced Qwen3.6-35B-A3B-singprobe, a streaming safety probe for Qwen3.6-35B-A3B

Ant Group 蚂蚁集团Chinese AIOpen weights
FACT STATEMENT

inclusionAI released Qwen3.6-35B-A3B-singprobe on Hugging Face under Apache-2.0. It is an intrinsic streaming guardrail probe built on Qwen/Qwen3.6-35B-A3B that scores query intent, response unsafety, and hallucination risk at every token with less than 0.5% decode-time overhead.

China context

Original name
inclusionAI
Outside China
Open weights · huggingface.co
Claims
Company-reported; not yet independently evaluated
For builders
Developers outside China can download the model from Hugging Face and integrate it via SGLang or vLLM branches to add streaming safety and hallucination detection to Qwen3.6-35B-A3B deployments.
For investors
The release of an open-weights guardrail probe by inclusionAI (Ant Group) indicates continued investment in safety tooling for open models, which may affect the competitive landscape for AI safety startups.
What happened

inclusionAI released Qwen3.6-35B-A3B-singprobe on Hugging Face under Apache-2.0. The model is an intrinsic streaming guardrail built on Qwen/Qwen3.6-35B-A3B that reuses the base model's hidden states during generation to score query intent, response unsafety, and hallucination risk at every token. It adds less than 0.5% decode-time overhead and has 4.2M probe parameters tapped at layers [12, 25, 38]. Evaluation results show F1 of 0.8699 on query intent classification, 0.8658 on response safety classification, R-AUC/T-AUC of 0.9889/0.9359 on streaming safety, and AUC of 0.8023 on hallucination detection. The model is supported through SGLang and vLLM integration branches.

Technical significance

The probe adds only 4.2M parameters and taps hidden states at layers [12, 25, 38] of the base model, achieving less than 0.5% decode-time overhead. It outputs a score dictionary per generated token with labels 0-9 covering 8 intents plus unsafe and hallucination. Evaluation results are company-reported and compare against reference baselines such as YuFeng-XGuard-Reason-8B, Qwen3Guard-Gen-8B-strict, Qwen3Guard-Stream-8B-strict, and DRIFT.

Industry impact

Developers using Qwen3.6-35B-A3B can add streaming safety and hallucination detection without deploying a separate safety model, reducing infrastructure cost and latency. This may pressure providers of standalone guardrail models to offer integrated alternatives.

Decision value

The model is open-weights under Apache-2.0, allowing commercial use and modification. It targets developers needing low-overhead safety monitoring for Qwen3.6-35B-A3B deployments.

What to watch

Adoption can be tracked by monitoring downloads and community usage of the Hugging Face model page, as well as integration activity in the SGLang and vLLM branches. Independent evaluation of the reported metrics would clarify performance relative to existing guardrails.

Latest in Chinese AI

  1. MiniMaxMiniMax open-sources MiniMax-Code-MiniApps repository for community-built plugins
  2. DeepSeekDeepSeek open-sources dsh-libreoffice-kit 0.1.0 for font-friendly Office conversion and rendering in Node.js
  3. DeepSeekDeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
  4. Shanghai AI LaboratoryShanghai AI Laboratory open-sources AdvancedMathBench for proof generation and verification
  5. Shanghai AI LaboratoryInternLM released a Qwen3-based model that grades mathematical proofs

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.