Event date · · MiniMaxAI

MiniMaxAI released MiniMax-Music3 on Hugging Face

MiniMax 稀宇科技Chinese AIOpen weights
FACT STATEMENT

MiniMaxAI released MiniMax-Music3 on Hugging Face. The model generates complete songs up to five minutes long from lyrics and a music description, producing 32 kHz, 16-bit stereo WAV audio. It uses an 8B Global LLM initialized from Qwen3-8B and a 0.6B Local LLM, with a Flow-VAE synthesis module adapted from MiniMax Speech. The training tokenizer uses eight layers of Residual Vector Quantization (RVQ), with the first semantic codebook containing 16,384 entries and the remaining seven acoustic codebooks containing 1,024 entries each. The model is supported by SGLang-Omni, diffusers, and ComfyUI.

China context

Original name
稀宇科技
Outside China
Open weights · huggingface.co
Claims
Company-reported; not yet independently evaluated
For builders
Developers outside China can access the model weights on Hugging Face and use supported inference frameworks like SGLang-Omni, diffusers, and ComfyUI to integrate music generation into applications.
For investors
The release of MiniMax-Music3 as open weights signals MiniMax's strategy to compete in the global AI music generation market and could attract developer mindshare and ecosystem growth.

Translated from Chinese. Quotes and facts link to the original sources.

What happened

MiniMaxAI released MiniMax-Music3, a high-performance music generation model, on Hugging Face. The model creates complete songs up to five minutes long conditioned on lyrics and a detailed music description. It combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE. The model produces 32 kHz, 16-bit stereo WAV audio and supports fine-grained control through structured captions and lyric section tags.

Technical significance

MiniMax-Music3 uses a hierarchical autoregressive architecture with a Global LLM (8B) initialized from Qwen3-8B to predict the first RVQ codebook frame by frame, and a Local LLM (0.6B) to predict the remaining acoustic codebooks. The synthesis module fuses the final hidden states of both LLMs, preserving richer acoustic information than discrete token decoding. The Flow-VAE architecture is adapted from MiniMax Speech and retrained for music. The training tokenizer uses eight layers of RVQ, with a 16,384-entry semantic codebook and seven 1,024-entry acoustic codebooks.

Industry impact

MiniMax-Music3 enters the competitive field of AI music generation, offering full-song generation up to five minutes with structural coherence and fine-grained control. The model's availability on Hugging Face and support for multiple inference frameworks (SGLang-Omni, diffusers, ComfyUI) lowers the barrier for developers and researchers to experiment with and build upon the model. The use of a hybrid LLM architecture and continuous hidden-state synthesis represents a notable technical approach in the open-weights music generation space.

Decision value

MiniMax-Music3 provides a tool for generating complete songs with expressive vocals and evolving arrangements, which could be valuable for content creators, game developers, and music producers. The model's open availability on Hugging Face may drive adoption and integration into creative workflows, while the underlying technology could be leveraged for other audio generation tasks.

What to watch

Observable next signals include the release of official evaluation benchmarks or comparisons with other music generation models, community adoption and fine-tuning efforts on Hugging Face, and the availability of the model through MiniMax's API platform. The model's performance on long-form music generation and its ability to follow structured captions will be key areas to watch.

CHINA AI WEEKLY

Get the week in Chinese AI, in English.

One weekly issue of verified model, company, robotics and policy changes, each with its original source and outside-China availability.