inclusionAI open-sourced Step-3.7-Flash-singprobe, a streaming safety probe for Step-3.7-Flash
inclusionAI released Step-3.7-Flash-singprobe, an Apache-2.0 licensed streaming guardrail probe built on stepfun-ai/Step-3.7-Flash, on Hugging Face. It adds less than 0.5% decode-time overhead and scores query intent, response unsafety, and hallucination risk at every token.
China context
- Original name
- inclusionAI
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the probe from Hugging Face and integrate it via SGLang or vLLM branches to add streaming safety scoring to Step-3.7-Flash deployments.
- For investors
- The release of a lightweight, Apache-2.0 safety probe by an Ant Group-affiliated entity may indicate a strategy to build an open ecosystem around StepFun models, potentially increasing adoption and competitive pressure on proprietary guardrail solutions.
inclusionAI released Step-3.7-Flash-singprobe on Hugging Face under Apache-2.0. The probe reuses the base model's hidden states during generation to score query intent, response unsafety, and hallucination risk at every token, adding less than 0.5% decode-time overhead. It has 8.13M parameters and taps layers [13, 28, 43]. Evaluation results are company-reported: query intent F1 0.8502, response safety F1 0.8555, streaming safety R-AUC 0.9858 / T-AUC 0.9295, hallucination detection AUC 0.7904. Benign-response false-positive rate is 0.05% average across 5 datasets. Training codes are available at inclusionAI/SingProbe.
The probe is an intrinsic guardrail that taps hidden states from layers [13, 28, 43] of the base model, avoiding a separate safety model. It outputs one score dictionary per generated token (labels 0–9). Deployment is supported through SGLang and vLLM integration branches, loading the probe by Hugging Face ID at server launch. The exact base-model/probe pair must be used: stepfun-ai/Step-3.7-Flash with this checkpoint.
Developers using Step-3.7-Flash can add streaming safety scoring with less than 0.5% decode overhead, reducing the need for separate guardrail models and lowering inference cost for safety-critical applications.
The Apache-2.0 license allows commercial use and modification, potentially lowering the barrier for enterprises to deploy streaming safety guardrails on Step-3.7-Flash without additional licensing costs.
Next signals to check: whether inclusionAI publishes the technical report with complete results, whether the SGLang/vLLM integration branches are merged upstream, and whether independent benchmarks confirm the reported F1/AUC scores.