NetEase Youdao released Confucius4-T3PO, a 14B simultaneous translation model, on Hugging Face
NetEase Youdao released Confucius4-T3PO, a 14-billion-parameter text-to-text simultaneous machine translation model, on Hugging Face under Apache-2.0. It supports streaming text input with adjustable latency modes and is built on Qwen2.5-14B-Instruct.
China context
- Original name
- 网易有道
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers can integrate a streaming simultaneous translation model with adjustable latency modes and Apache-2.0 licensing, enabling custom real-time translation applications without vendor lock-in.
- For investors
- NetEase Youdao's release of an open-weights simultaneous translation model signals a strategic move to capture developer mindshare in the real-time translation market, pressuring commercial API providers.
NetEase Youdao released Confucius4-T3PO, a 14-billion-parameter text-to-text simultaneous machine translation (SiMT) model, on Hugging Face under the Apache-2.0 license. The model supports streaming text input and real-time READ/WRITE decisions, with adjustable latency modes for different quality-latency trade-offs. It is built on Qwen2.5-14B-Instruct and retains general instruction-following ability. The release includes a GGUF version and a locally deployable Web UI and streaming client.
Confucius4-T3PO uses a three-stage pipeline: high-quality segment-aligned data construction, streaming cold-start, and Pareto-aware reinforcement learning for joint quality-latency optimization. It employs an interleaved history protocol for KV-cache reuse and append-only committed translations. The model supports character- and word-level chunk input and can be cascaded with an external streaming ASR model for speech-to-text translation.
Developers outside China can now use an open-weights simultaneous translation model with Apache-2.0 licensing, reducing the cost and complexity of building low-latency translation services compared to commercial APIs. Competitors offering real-time translation must now match or exceed its adjustable latency-quality tiers or risk losing developer mindshare.
The model enables developers to deploy simultaneous translation services without per-token API costs, reducing operational expenses for real-time translation applications. Its Apache-2.0 license allows commercial use, and the availability of GGUF and local deployment tools lowers infrastructure barriers.
The team plans to release a technical report with further details on training method, data, and implementation. Evaluation on public benchmarks for Chinese-English and English-Chinese simultaneous translation is provided, with comparison to InfiniSST, EAST, and two commercial systems. Cross-lingual generalization to Japanese-Chinese is noted but not rigorously evaluated.