inclusionAI open-sourced SingProbe, a streaming safety probe for Qwen3.5-35B-A3B
inclusionAI released Qwen3.5-35B-A3B-singprobe on Hugging Face under Apache-2.0. It is an intrinsic streaming guardrail probe built on Qwen/Qwen3.5-35B-A3B that scores query intent, response unsafety, and hallucination risk at every token with less than 0.5% decode-time overhead.
China context
- Original name
- inclusionAI/Qwen3.5-35B-A3B-singprobe
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the weights from Hugging Face and integrate the probe via SGLang or vLLM branches to add streaming safety and hallucination detection to Qwen3.5-35B-A3B deployments.
- For investors
- The release of a low-overhead safety probe by an Ant Group-affiliated entity may indicate a focus on production-grade safety tooling for open-weights models, which could affect the competitive landscape for AI safety startups.
inclusionAI released Qwen3.5-35B-A3B-singprobe on Hugging Face under Apache-2.0. The model is an intrinsic streaming guardrail probe built on Qwen/Qwen3.5-35B-A3B that reuses the base model's hidden states during generation to score query intent, response unsafety, and hallucination risk at every token. It adds less than 0.5% decode-time overhead and has 4.2M probe parameters tapped at layers [12, 25, 38]. Evaluation results show F1 of 0.8751 for query intent classification, 0.8672 for response safety classification, R-AUC/T-AUC of 0.9902/0.9374 for streaming safety, and AUC of 0.8117 for hallucination detection. The benign-response false-positive rate is 0.03% average across 5 datasets. The model is supported through SGLang and vLLM integration branches.
SingProbe is a lightweight probe that reuses the base model's hidden states during generation, adding less than 0.5% decode-time overhead. It has 4.2M probe parameters tapped at layers [12, 25, 38] and outputs 8 intents plus unsafe and hallucination scores per token. Evaluation results are company-reported and include F1 of 0.8751 for query intent classification, 0.8672 for response safety classification, R-AUC/T-AUC of 0.9902/0.9374 for streaming safety, and AUC of 0.8117 for hallucination detection. The benign-response false-positive rate is 0.03% average across 5 datasets.
Developers using Qwen3.5-35B-A3B can add streaming safety and hallucination detection without running a separate safety model, reducing inference cost and complexity. This may pressure providers of standalone guardrail models to offer integrated or lower-overhead alternatives.
The probe offers a low-overhead way to add safety and hallucination detection to Qwen3.5-35B-A3B deployments, potentially reducing the need for separate safety models and lowering operational costs for developers.
Independent evaluation of the reported metrics and real-world false-positive rates on diverse traffic is needed. Adoption can be tracked via Hugging Face downloads, GitHub stars on the integration branches, and community reports of production use.