inclusionAI open-sourced Qwen3.6-35B-A3B-singprobe, a streaming safety probe for Qwen3.6-35B-A3B
inclusionAI released Qwen3.6-35B-A3B-singprobe on Hugging Face under Apache-2.0. It is an intrinsic streaming guardrail probe built on Qwen/Qwen3.6-35B-A3B that scores query intent, response unsafety, and hallucination risk at every token with less than 0.5% decode-time overhead.
China context
- Original name
- inclusionAI
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the model from Hugging Face and integrate it via SGLang or vLLM branches to add streaming safety and hallucination detection to Qwen3.6-35B-A3B deployments.
- For investors
- The release of an open-weights guardrail probe by inclusionAI (Ant Group) indicates continued investment in safety tooling for open models, which may affect the competitive landscape for AI safety startups.
inclusionAI released Qwen3.6-35B-A3B-singprobe on Hugging Face under Apache-2.0. The model is an intrinsic streaming guardrail built on Qwen/Qwen3.6-35B-A3B that reuses the base model's hidden states during generation to score query intent, response unsafety, and hallucination risk at every token. It adds less than 0.5% decode-time overhead and has 4.2M probe parameters tapped at layers [12, 25, 38]. Evaluation results show F1 of 0.8699 on query intent classification, 0.8658 on response safety classification, R-AUC/T-AUC of 0.9889/0.9359 on streaming safety, and AUC of 0.8023 on hallucination detection. The model is supported through SGLang and vLLM integration branches.
The probe adds only 4.2M parameters and taps hidden states at layers [12, 25, 38] of the base model, achieving less than 0.5% decode-time overhead. It outputs a score dictionary per generated token with labels 0-9 covering 8 intents plus unsafe and hallucination. Evaluation results are company-reported and compare against reference baselines such as YuFeng-XGuard-Reason-8B, Qwen3Guard-Gen-8B-strict, Qwen3Guard-Stream-8B-strict, and DRIFT.
Developers using Qwen3.6-35B-A3B can add streaming safety and hallucination detection without deploying a separate safety model, reducing infrastructure cost and latency. This may pressure providers of standalone guardrail models to offer integrated alternatives.
The model is open-weights under Apache-2.0, allowing commercial use and modification. It targets developers needing low-overhead safety monitoring for Qwen3.6-35B-A3B deployments.
Adoption can be tracked by monitoring downloads and community usage of the Hugging Face model page, as well as integration activity in the SGLang and vLLM branches. Independent evaluation of the reported metrics would clarify performance relative to existing guardrails.