Event date · · Zhipu AI

Zhipu AI released GLM-5.3-Flash, a 320B-parameter multimodal model with 18B active parameters, on Hugging Face

Zhipu AI 智谱Chinese AIOpen weights
FACT STATEMENT

Zhipu AI (via zai-org) released GLM-5.3-Flash on Hugging Face under the MIT license. The model has 320B total parameters and 18B active parameters, is natively multimodal (image-text-to-text), and is the first in the GLM-5 series. It outperforms GLM-5.2 at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It uses a hybrid sparse and linear attention architecture and Manifold-Constrained Hyper-Connections (mHC), trained on a 30T-token multimodal corpus. Deployment is supported via SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth. The reasoning effort parameter accepts low, high, and max (default max).

China context

Original name
智谱
Outside China
Open weights · huggingface.co
Claims
Company-reported; not yet independently evaluated
For builders
Developers outside China can download the MIT-licensed weights from Hugging Face and deploy using SGLang, vLLM, Transformers, or other supported frameworks. The model's sparse architecture may reduce serving costs for long-context applications.
For investors
The release of a high-efficiency multimodal model with competitive coding/agentic claims at one-tenth the price of its predecessor may pressure pricing and adoption of Western models. Investors should monitor independent benchmark results and API pricing on Z.ai.

Translated from Chinese. Quotes and facts link to the original sources.

What happened

Zhipu AI released GLM-5.3-Flash, a natively multimodal model with 320B total and 18B active parameters, under MIT license on Hugging Face. It outperforms GLM-5.2 at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. The model introduces a hybrid sparse and linear attention architecture and mHC, trained on a 30T-token multimodal corpus. It supports deployment via multiple frameworks and offers a reasoning effort parameter.

Technical significance

GLM-5.3-Flash uses a hybrid sparse and linear attention mechanism to reduce long-context serving costs while maintaining precision. It also employs Manifold-Constrained Hyper-Connections (mHC) for scaling efficiency. The model is natively multimodal, handling image and text inputs. It supports a reasoning effort parameter with three levels (low, high, max) to control thinking budget.

Industry impact

The release of GLM-5.3-Flash with 18B active parameters out of 320B total parameters signals a trend toward sparse architectures for cost-efficient inference. The MIT license and support for multiple serving frameworks (SGLang, vLLM, etc.) lower barriers for adoption. The claim of approaching Claude Opus 4.8 on coding and agentic benchmarks positions it competitively against Western frontier models.

Decision value

For developers, GLM-5.3-Flash offers a permissively licensed multimodal model with efficient inference and flexible deployment. For enterprises, the one-tenth price compared to GLM-5.2 and competitive coding/agentic performance may reduce costs for AI applications. The model's availability on Hugging Face and multiple serving frameworks facilitates integration.

What to watch

Observable next signals include independent benchmark evaluations (e.g., Artificial Analysis for GDPval-AA v2), adoption metrics on Hugging Face (downloads, likes), and API pricing/availability on Z.ai. The release of the full GLM-5.3 model (text-generation) may follow.

CHINA AI WEEKLY

Get the week in Chinese AI, in English.

One weekly issue of verified model, company, robotics and policy changes, each with its original source and outside-China availability.