inclusionAI open-sourced Qwen3-8B-singprobe, a streaming safety and hallucination probe
inclusionAI released Qwen3-8B-singprobe, an Apache-2.0 licensed streaming guardrail probe built on Qwen3-8B, on Hugging Face. It adds less than 0.5% decode-time overhead and scores query intent, response unsafety, and hallucination risk at every token.
China context
- Original name
- inclusionAI
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the Apache-2.0 weights from Hugging Face and integrate via SGLang or vLLM branches to add streaming safety scoring to Qwen3-8B.
- For investors
- The release signals Ant Group's continued investment in open-weights safety tooling, which may pressure commercial guardrail vendors.
inclusionAI released Qwen3-8B-singprobe on Hugging Face under Apache-2.0. The probe reuses Qwen3-8B hidden states to score query intent, response unsafety, and hallucination risk per token, with less than 0.5% decode-time overhead. It is supported via SGLang and vLLM integration branches.
The probe taps layers [10, 22, 34] of Qwen3-8B and adds 8.13M parameters. Reported F1 scores: 0.8687 for query intent, 0.8616 for response safety, 0.9878 R-AUC for streaming safety, and 0.7849 AUC for hallucination detection. Benign-response false-positive rate is 0.03% across 5 datasets.
Developers deploying Qwen3-8B can add streaming safety and hallucination scoring without a separate safety model, reducing inference cost and latency. Competitors offering guardrail models must match sub-0.5% overhead and per-token scoring to remain attractive.
The probe lowers the cost and complexity of adding safety and hallucination detection to Qwen3-8B deployments, potentially increasing enterprise adoption of the base model.
Next signals to check: adoption of the SGLang/vLLM integration branches, independent benchmarks against the reported metrics, and whether inclusionAI releases probes for other base models.