inclusionAI open-sourced Qwen3.5-2B-singprobe, a streaming safety probe for Qwen3.5-2B
inclusionAI released Qwen3.5-2B-singprobe, an Apache-2.0 licensed streaming guardrail probe built on Qwen/Qwen3.5-2B, on Hugging Face. It adds less than 0.5% decode-time overhead and scores query intent, response unsafety, and hallucination risk at every token.
China context
- Original name
- inclusionAI/Qwen3.5-2B-singprobe
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the weights from Hugging Face and integrate via SGLang or vLLM branches to add streaming safety checks to Qwen3.5-2B with minimal overhead.
- For investors
- The release of a low-overhead safety probe by an Ant Group-affiliated entity signals continued investment in open-weights safety tooling, which may affect the competitive landscape for commercial guardrail APIs.
inclusionAI released Qwen3.5-2B-singprobe on Hugging Face under Apache-2.0. The model is an intrinsic streaming guardrail that reuses the base model's hidden states during generation to score query intent, response unsafety, and hallucination risk at every token, with less than 0.5% decode-time overhead. It has 4.2M probe parameters tapped from layers [6, 14, 22] and outputs 8 intents plus unsafe and hallucination scores. Evaluation results show F1 of 0.8542 on query intent classification, 0.8516 on response safety classification, R-AUC/T-AUC of 0.9824/0.9243 on streaming safety, and AUC of 0.7642 on hallucination detection. The model is supported through SGLang and vLLM integration branches.
The probe adds only 4.2M parameters and taps hidden states from layers [6, 14, 22] of Qwen3.5-2B, achieving less than 0.5% decode-time overhead. It outputs a score dictionary per generated token with labels 0-9. The model card reports a benign-response false-positive rate of 0.08% average across 5 datasets. Training codes are available at inclusionAI/SingProbe.
Developers using Qwen3.5-2B can add streaming safety and hallucination detection without deploying a separate safety model, reducing inference cost and latency compared to external guardrails. This may pressure providers of standalone guardrail models to offer tighter integration or lower overhead.
The model is open-weights under Apache-2.0, allowing commercial use and modification. It targets developers needing low-overhead safety filtering for Qwen3.5-2B deployments.
Adoption can be tracked via Hugging Face downloads and community usage of the SGLang/vLLM integration branches. Independent evaluation of the reported metrics and false-positive rate would verify performance claims.