ModelBest released JustRL-II-base-model on Hugging Face
ModelBest released JustRL-II-base-model on Hugging Face, an RL initialization checkpoint for long chain-of-thought reasoning. It scores about 61% on AIME 2025 before RL, and the full JustRL II recipe reaches 81% in ~300 RL steps.
China context
- Original name
- 面壁智能
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the checkpoint from Hugging Face and use it to reproduce the JustRL II recipe or as a starting point for long-CoT RL research.
- For investors
- The release of a reproducible RL checkpoint by ModelBest indicates active open-weights competition in the small-model reasoning space, which may pressure other labs to release similar artifacts.
ModelBest (OpenBMB) released JustRL-II-base-model on Hugging Face. It is the RL initialization checkpoint used in the JustRL II blog, which scales small LLMs to 128K reasoning with a critic. The checkpoint is a small language model trained to produce long chain-of-thought responses inside <think> tags. Before RL, it scores about 61% on AIME 2025 under the blog's evaluation protocol. The full JustRL II recipe reaches 81% on AIME 2025 in ~300 RL steps, while a standard GRPO baseline plateaus around 74%. The release includes bf16 weights, Llama-style config, tokenizer, and chat template with thinking-mode support. It is intended for reproducing the JustRL II recipe and research on long-CoT RL for small models.
The checkpoint uses a LlamaForCausalLM architecture with max position embeddings of 65536, but the RL runs use a 128k-token generation budget. It has two end-of-sequence token ids ([1, 130073]) that must be passed to generation calls. The model emits reasoning inside <think>...</think> before the final answer. The JustRL II recipe uses a critic-equipped GRPO with a learned value model, GAE with length-adaptive λ, and tail-only overlong control. The data pipeline audits ~100k math problems and leaves 32,412 problems after removing those the checkpoint solves 8/8.
ModelBest's release of the JustRL II base checkpoint gives researchers outside China a reproducible starting point for long-CoT RL on small models, lowering the cost of replicating the 81% AIME 2025 result. Competitors can now benchmark their own RL recipes against this exact initialization.
The release enables developers to reproduce the JustRL II recipe and ablations from the same initialization, potentially reducing R&D costs for long-CoT RL on small models. It also provides a benchmark for evaluating RL algorithms on mathematical reasoning.
The next verifiable signal is whether ModelBest releases the post-RL JustRL II model weights or the full training code. Another signal is independent reproduction of the 81% AIME 2025 score using the released checkpoint and recipe.