Tencent released Simple-Attention-Sparsification, a method for efficient sparse attention in Qwen3 models
Tencent released Simple-Attention-Sparsification (SAS), a method that learns to rank and select KV blocks for each query using continuous gates, with checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B. The checkpoints are router-only and require the corresponding Qwen3 base model and a sparse SGLang backend.
China context
- Original name
- 腾讯
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can access the checkpoints on Hugging Face, but must use Tencent's sglang-blocksparse fork and the corresponding Qwen3 base model, which may require additional setup.
- For investors
- The release indicates Tencent's continued investment in efficient inference methods, but the router-only nature and custom backend may limit near-term commercial impact.
Tencent released Simple-Attention-Sparsification (SAS) on Hugging Face. SAS learns to rank and select KV blocks for each query. Unlike methods that train a sparse-attention selector by distilling dense attention scores, SAS adds continuous gates to the selected blocks so that the language-modeling loss can optimize context ranking end to end. Released checkpoints include Qwen3-4B-AttnGates, Qwen3-8B-AttnGates, and Qwen3-14B-AttnGates. Each directory is an SGLang-compatible AttnGates package containing attn gate weights, config.json, and tokenizer/chat-template files. The checkpoints are router-only, not standalone language models; the Qwen3 backbone was frozen during training and is not included. Inference requires the corresponding Qwen3 base model and the seer_attn backend in Tencent's sglang-blocksparse fork. All three checkpoints use the same sparse-attention setup: KV block size 64 tokens, training Top-K 31 historical blocks, gate hidden size 128, query projection Qproj, key block pooling max + min + average, Q/K normalization enabled, gate RoPE enabled, training sequence length 32,768 tokens, and training data OpenR1-Math-220k. The released setup uses a 2,048-token sparse decode budget by default, and the same checkpoints can be evaluated with 1,024-, 2,048-, or 4,096-token budgets without retraining.
SAS introduces continuous gates on selected KV blocks, providing a differentiable path through discrete Top-K block selection. This allows end-to-end optimization of context ranking with the language modeling loss, unlike distillation-based sparse attention. The checkpoints are router-only, with the base model frozen, and require a custom SGLang backend (seer_attn) for inference.
Developers using Qwen3 models can reduce attention memory and compute by adopting SAS, but they must integrate Tencent's sglang-blocksparse fork and accept router-only checkpoints that depend on the original base model.
SAS offers a potential efficiency improvement for long-context inference on Qwen3 models, which could lower serving costs for developers who adopt the custom backend. However, the requirement to use a specific fork and the lack of standalone checkpoints may limit immediate adoption.
Observable next signals include whether Tencent releases training code or full model weights, whether the method is adopted by other model providers, and whether independent benchmarks confirm efficiency gains on standard tasks.