NetEase Youdao released Confucius4-R2T2, a low-latency streaming ASR model, on Hugging Face
NetEase Youdao released Confucius4-R2T2, a low-latency streaming automatic speech recognition model, on Hugging Face. The model achieves 200–600 ms average latency with accuracy close to offline recognition and supports configurable decoding chunks from 80 ms to 2 s.
China context
- Original name
- 网易有道
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can access the model weights on Hugging Face and integrate streaming ASR with append-only output into real-time applications.
- For investors
- The release of a competitive open-weights streaming ASR model by NetEase Youdao may pressure commercial ASR API pricing and accelerate adoption of self-hosted speech recognition.
Confucius4-R2T2 is a streaming ASR model built on Qwen3-ASR, featuring append-only output that commits transcript text permanently without revising previous words. It is optimized for Chinese and English, supports additional languages, and provides vLLM and Hugging Face transformers backends. The model is available on Hugging Face and ModelScope under a NetEase model license, with code under Apache 2.0.
The model uses a Longest Stable Prefix (LSP) learning paradigm with stable-prefix data, forced time-alignment data, and token-level audio segmentation to dynamically determine when a stable prefix can be emitted. It achieves state-of-the-art latency and recognition quality among open-source models while remaining competitive with closed-source systems, with no loss in offline accuracy.
Developers building real-time transcription, live captioning, or LLM agent pipelines can now use an open-weights model that commits text permanently, eliminating disruptive revisions and visual flickering in streaming applications. This reduces integration complexity for low-latency ASR compared to models that revise previous outputs.
The model's append-only output and low latency make it suitable for real-time applications such as live captioning, simultaneous translation, and downstream NLP pipelines. Its availability on Hugging Face and ModelScope with a permissive code license may accelerate adoption among developers.
A technical report on the LSP learning paradigm is expected to be released soon. Verification of the claimed state-of-the-art performance against independent benchmarks and adoption in production systems will be key signals to watch.