Qwen released Qwen3.8-2.4T-A95B on Hugging Face
Qwen (Alibaba) released Qwen3.8-2.4T-A95B on Hugging Face. The model has 2.4T total parameters, 95B activated, 92 layers, 512 experts (10 routed + 1 shared), context length 262,144 natively extensible to 1,010,000 tokens. It is a causal language model with pre-training and post-training. Qwen3.8-Max is the official API version with vision input, non-thinking support, 1M context length by default, and built-in tools.
China context
- Original name
- Qwen3.8-2.4T-A95B
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the weights from Hugging Face and run the model with vLLM, SGLang, or TokenSpeed.
- For investors
- The open release of a Qwen-Max-class model may pressure other frontier labs to release open weights or adjust API pricing.
Qwen released Qwen3.8-2.4T-A95B, a 2.4T-parameter open-weights model with 95B activated parameters, on Hugging Face. The model is compatible with vLLM, SGLang, and TokenSpeed. Qwen3.8-Max, the API version, adds vision input, non-thinking support, 1M context length, and built-in tools.
Qwen3.8-2.4T-A95B uses a hybrid architecture with 23 blocks of 3 Gated DeltaNet layers followed by 1 Gated Attention layer, each with MoE. It has 512 experts with 10 routed and 1 shared per token. The model supports multi-token prediction and reasoning effort control.
This is the first Qwen-Max-class model released as open weights, following community adoption of Qwen3.5 and Qwen3.6. The release includes a managed API option via Qwen Cloud.
Open-weights release of a Qwen-Max-class model allows developers to self-host a frontier-scale model, while Qwen Cloud offers a managed API with additional features.
Observable next signals include adoption of Qwen3.8-2.4T-A95B in downstream tools and benchmarks, and whether Qwen3.8-Max API pricing or availability changes.