Tencent Hunyuan open-sources SAS sparse attention training pipeline
Tencent Hunyuan released the Simple-Attention-Sparsification (SAS) repository with training recipes for Qwen3-4B, 8B, and 14B models, trainable on 8 NVIDIA H20 GPUs. The code includes evaluation scripts and an SGLang-based sparse attention backend.
China context
- Original name
- 腾讯混元
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- The repository includes training recipes and an SGLang fork, enabling developers to implement block-sparse attention in their own models.
- For investors
- Tencent's open-sourcing of efficient attention methods may pressure competitors to release similar efficiency improvements, affecting the cost structure of AI inference.
Tencent Hunyuan published the SAS repository, implementing a gated sparse-attention mechanism that learns per-query block selection end-to-end with the language modeling loss. It provides training recipes for Qwen3-4B, 8B, and 14B, requiring 8 NVIDIA H20 GPUs, and includes evaluation scripts for reasoning, long-context, and agentic benchmarks, plus an SGLang fork for block-sparse inference.
SAS replaces discrete Top-K selection with continuous soft gates on selected blocks, creating a differentiable path from the LM loss to the selector. This avoids auxiliary distillation objectives. The repository includes a custom SGLang fork (sglang-blocksparse) for efficient block-sparse decoding and supports benchmarks like MATH, GPQA-Diamond, LongBench-E, BFCL, and VitaBench.
Developers outside China can now reproduce and extend Tencent's sparse attention method, reducing inference compute costs for long-context models without relying on proprietary distillation pipelines. The open-sourced training recipes and SGLang backend lower the barrier to adopting block-sparse attention in production systems.
The release provides a cost-effective path to sparse attention for organizations with limited GPU resources, as training requires only 8 H20 GPUs. The open-source SGLang backend could reduce serving costs for long-context applications.
Watch for independent benchmarks comparing SAS against dense baselines on the provided tasks, and for adoption of the sglang-blocksparse fork in other projects. The paper's arXiv link (2609.13141) may provide further details on performance and limitations.