BAAI released AREX-2, a 27B self-improving agent model, on Hugging Face
BAAI released AREX-2, a 27B-parameter long-horizon agent model, on Hugging Face. It learns to improve solutions over multiple test-time rounds via propose, measure, reflect, and revise. The model is trained on machine-learning and algorithmic-programming tasks with verifiable feedback.
China context
- Original name
- 北京智源人工智能研究院
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Builders can download the weights from Hugging Face and experiment with a 27B agent model that uses test-time self-improvement, reducing the need for larger models in coding and research agents if the reported benchmark scores are confirmed.
- For investors
- Investors should monitor whether AREX-2's self-improvement approach gains traction among developers, as it could shift demand toward smaller, more efficient open-weight models for agentic workloads.
AREX-2 is a 27B-parameter dense Qwen3.8-compatible multimodal model with a 262,144-token context length. It is designed for long-horizon self-improvement, using feedback-driven reflection to refine solutions across coding, machine-learning engineering, deep research, and general agentic reasoning. The model card reports Frontier-CS score of 70.7 and MLE-Lite score of 81.8, outperforming several larger open-weight models on these benchmarks.
AREX-2 uses a propose-measure-reflect-revise loop to iteratively improve solutions during test time. It is trained on tasks with verifiable feedback, and the learned self-improvement transfers to deep research without additional search trajectories. The model is dense and multimodal, compatible with Qwen3.8 architecture.
Developers outside China can now access a 27B open-weight agent model that claims to outperform larger models like DeepSeek-V4-Pro on MLE-Lite, reducing the compute needed for agentic coding and research tasks if the reported scores hold in independent tests.
AREX-2 offers a smaller, open-weight alternative for agentic tasks, which could lower inference costs for developers building self-improving agents.
Observable next signals include independent evaluations of AREX-2 on Frontier-CS and MLE-Lite, and whether the self-improvement loop is adopted in downstream agent frameworks.