Event date · · Tencent

Tencent Hunyuan open-sources SAS sparse attention training pipeline

Tencent 腾讯Chinese AIOpen weights
FACT STATEMENT

Tencent Hunyuan released the Simple-Attention-Sparsification (SAS) repository with training recipes for Qwen3-4B, 8B, and 14B models, trainable on 8 NVIDIA H20 GPUs. The code includes evaluation scripts and an SGLang-based sparse attention backend.

China context

Original name
腾讯混元
Outside China
Open weights · github.com
Claims
Company-reported; not yet independently evaluated
For builders
The repository includes training recipes and an SGLang fork, enabling developers to implement block-sparse attention in their own models.
For investors
Tencent's open-sourcing of efficient attention methods may pressure competitors to release similar efficiency improvements, affecting the cost structure of AI inference.
What happened

Tencent Hunyuan published the SAS repository, implementing a gated sparse-attention mechanism that learns per-query block selection end-to-end with the language modeling loss. It provides training recipes for Qwen3-4B, 8B, and 14B, requiring 8 NVIDIA H20 GPUs, and includes evaluation scripts for reasoning, long-context, and agentic benchmarks, plus an SGLang fork for block-sparse inference.

Technical significance

SAS replaces discrete Top-K selection with continuous soft gates on selected blocks, creating a differentiable path from the LM loss to the selector. This avoids auxiliary distillation objectives. The repository includes a custom SGLang fork (sglang-blocksparse) for efficient block-sparse decoding and supports benchmarks like MATH, GPQA-Diamond, LongBench-E, BFCL, and VitaBench.

Industry impact

Developers outside China can now reproduce and extend Tencent's sparse attention method, reducing inference compute costs for long-context models without relying on proprietary distillation pipelines. The open-sourced training recipes and SGLang backend lower the barrier to adopting block-sparse attention in production systems.

Decision value

The release provides a cost-effective path to sparse attention for organizations with limited GPU resources, as training requires only 8 H20 GPUs. The open-source SGLang backend could reduce serving costs for long-context applications.

What to watch

Watch for independent benchmarks comparing SAS against dense baselines on the provided tasks, and for adoption of the sglang-blocksparse fork in other projects. The paper's arXiv link (2609.13141) may provide further details on performance and limitations.

Latest in Chinese AI

  1. MiniMaxMiniMax open-sources MiniMax-Code-MiniApps repository for community-built plugins
  2. DeepSeekDeepSeek open-sources dsh-libreoffice-kit 0.1.0 for font-friendly Office conversion and rendering in Node.js
  3. DeepSeekDeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
  4. Shanghai AI LaboratoryShanghai AI Laboratory open-sources AdvancedMathBench for proof generation and verification
  5. Shanghai AI LaboratoryInternLM released a Qwen3-based model that grades mathematical proofs

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.