Reportedly: Moonshot AI releases k0-math, a math model it says beats OpenAI o1 on four benchmarks
Reported by GeekPark · not yet confirmed by the company or a second independent outlet. We update this page when it is.
Moonshot AI released k0-math, a reasoning-reinforced math model, on November 16, 2024. The company reported it outperforms OpenAI o1-mini and o1-preview on four math benchmarks: middle school, high school, graduate entrance exams, and MATH. Kimi assistant had over 36 million monthly active users in October 2024.
China context
- Original name
- 月之暗面
- Outside China
- Not stated in the sources yet
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China should verify k0-math's benchmark claims independently before adopting it, as the reported performance is company-reported and not yet confirmed by third-party evaluations.
- For investors
- Investors should monitor whether Moonshot AI's focus on retention and growth translates into sustained user engagement, as the company reported over 36 million MAU but did not disclose retention metrics.
Translated from Chinese. Quotes and facts link to the original sources.
On the first anniversary of Kimi's full launch, Moonshot AI introduced k0-math, a math model built with reinforcement learning, synthetic data, and chain-of-thought. The company reported k0-math surpasses OpenAI o1-mini and o1-preview on four math benchmarks: middle school, high school, graduate entrance exams, and MATH. Founder Yang Zhilin said the model handles most difficult math problems but over-thinks simple ones and cannot yet solve geometry problems that are hard to express in LaTeX. Kimi assistant had over 36 million monthly active users in October 2024. The new model and Kimi Explorer features will roll out on web and app.
k0-math uses reinforcement learning and chain-of-thought without length constraints, which the company says causes over-thinking on simple problems like '1+1'. The model cannot handle geometry problems that are difficult to express in LaTeX. These limitations are company-reported and not independently verified.
Chinese AI labs are now directly targeting OpenAI's o1 reasoning models with specialized math models, which could shift developer choices for math-heavy applications if k0-math's benchmark claims hold in independent tests.
For developers building math-intensive applications, k0-math offers a potential alternative to OpenAI o1 if its performance is confirmed. For investors, Moonshot AI's focus on retention and growth, with over 36 million MAU, indicates a shift toward product stickiness rather than pure model releases.
The next verifiable signal is whether k0-math appears on public leaderboards or is independently benchmarked against o1. Also watch for the rollout of k0-math and Kimi Explorer features on web and app, and any change in Kimi's monthly active user numbers.