DeepSeek-V3 Open Source: Training Efficiency Becomes a New Variable in Global Model Competition
DeepSeek released and open-sourced DeepSeek-V3 weights and technical report, featuring 671B MoE, 37B activated parameters, and FP8 training.
DeepSeek uses architecture, system, and engineering synergy to lower frontier model training costs, demonstrating that different resource conditions can form independent frontier innovation paths.
MLA, DeepSeekMoE, auxiliary-loss-free load balancing, MTP, and FP8 jointly improve training and inference efficiency.
Model competition adds a core dimension of compute output per unit, and drives adaptation of domestic hardware and inference stacks.
Cost efficiency will compress general API gross margins, but expand the space for models to enter low-unit-price businesses.
Observe independent reproduction, real API load, domestic compute adaptation, and subsequent inference models.