inclusionAI open-sourced SingProbe, a streaming guardrail for gpt-oss-120b
inclusionAI released gpt-oss-120b-singprobe, an Apache-2.0 licensed streaming guardrail probe built on openai/gpt-oss-120b, adding less than 0.5% decode-time overhead. It scores query intent, response unsafety, and hallucination risk at every token using 5.8M probe parameters tapped from layers [10, 22, 34].
China context
- Original name
- inclusionAI
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the probe from Hugging Face and integrate it with SGLang or vLLM to add streaming safety scoring to openai/gpt-oss-120b deployments.
- For investors
- The release of an open-weights safety probe by inclusionAI (Ant Group) may indicate a strategic focus on AI safety tooling, which could influence the competitive landscape for guardrail solutions.
inclusionAI released gpt-oss-120b-singprobe on Hugging Face under Apache-2.0. The probe reuses hidden states from openai/gpt-oss-120b during generation to score query intent, response unsafety, and hallucination risk at every token, with less than 0.5% decode-time overhead. Evaluation results include F1 0.8699 for query intent classification, F1 0.8652 for response safety classification, R-AUC/T-AUC 0.9881/0.9454 for streaming safety, and AUC 0.7815 for hallucination detection. The model card provides quick start instructions via SGLang and vLLM integration branches.
SingProbe is an intrinsic streaming guardrail that reuses the base model's hidden states from layers [10, 22, 34] with 5.8M probe parameters, avoiding a separate safety model. It outputs a score dictionary per generated token (labels 0-9) and reports a benign-response false-positive rate of 0.06% average across 5 datasets. The probe must be paired with the exact base model openai/gpt-oss-120b.
Developers using openai/gpt-oss-120b can add streaming safety scoring with minimal overhead, potentially reducing the need for separate guardrail models and lowering inference costs for safety-critical applications.
The release provides a lightweight, open-source safety mechanism for large language models, which may appeal to enterprises needing real-time content moderation without significant latency or cost increases.
Adoption can be checked by monitoring Hugging Face downloads, community forks, or integration into SGLang/vLLM main branches. Independent evaluation of the reported metrics would confirm performance outside the company's own benchmarks.