inclusionAI open-sourced Qwen3.6-27B-singprobe, a streaming safety probe for Qwen3.6-27B
inclusionAI released Qwen3.6-27B-singprobe, an Apache-2.0 licensed streaming guardrail probe built on Qwen/Qwen3.6-27B, on Hugging Face. It adds less than 0.5% decode-time overhead and scores query intent, response unsafety, and hallucination risk at every token.
China context
- Original name
- inclusionAI
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the open-weights model from Hugging Face and integrate it via SGLang or vLLM to add streaming safety scoring to Qwen3.6-27B deployments.
- For investors
- The release of a lightweight, open-weights safety probe by an Ant Group-affiliated entity may pressure commercial guardrail providers by offering a low-cost alternative.
inclusionAI released Qwen3.6-27B-singprobe on Hugging Face under Apache-2.0. The model is an intrinsic streaming guardrail that reuses the base model's hidden states to score query intent, response unsafety, and hallucination risk at every token, adding less than 0.5% decode-time overhead. It has 10.1M probe parameters tapped from layers [20, 41, 62] and outputs 8 intents plus unsafe and hallucination scores. Evaluation results show F1 of 0.8718 on query intent classification, 0.8692 on response safety classification, R-AUC/T-AUC of 0.9864/0.9270 on streaming safety, and AUC of 0.8094 on hallucination detection. The model is supported through SGLang and vLLM integration branches.
The probe adds only 10.1M parameters and taps three layers of the 27B base model, achieving less than 0.5% decode-time overhead. It outputs a score dictionary per generated token with labels 0-9. The model card specifies using the exact base-model/probe pair Qwen/Qwen3.6-27B with this checkpoint. Training codes are available at inclusionAI/SingProbe.
Developers using Qwen3.6-27B can add streaming safety and hallucination detection without deploying a separate safety model, reducing infrastructure cost and latency. This positions inclusionAI's probe as a lightweight alternative to standalone guard models like Qwen3Guard.
The Apache-2.0 license allows commercial use, and the low overhead makes it attractive for production deployments requiring real-time safety scoring. The probe's performance is competitive with dedicated guard models, potentially reducing the need for separate safety infrastructure.
Adoption can be tracked via downloads and community usage on the Hugging Face model page, and by monitoring integration activity in the SGLang and vLLM branches. The technical report and training code release may lead to independent evaluations.