A paper titled 'Training AI Scientists to Replicate Research' was published on arXiv (cs.AI) on 2026-08-13. It introduces Replica, a scalable task space for paper replication, and an auto-generated rubric-based judge. The authors post-train Faraday, a 27B-parameter AI Scientist agent, which surpasses Claude Opus 4.8 and GPT-5.5 on held-out replication tasks.
The paper addresses paper replicability as a cornerstone of scientific knowledge. It develops Replica, a scalable task space for paper replication, and introduces an auto-generated rubric-based judge with low noise that agrees with human assessment. Faraday, a 27B-parameter AI Scientist agent, is post-trained using coding agents as tools and surpasses Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis shows Faraday adopts a more scientifically-principled approach. The authors believe this provides a stepping stone towards AI agents capable of long-horizon scientific innovation without complex harnesses.
The approach combines a scalable task space (Replica) with an auto-generated rubric-based judge for reward signal, enabling post-training of a 27B-parameter agent (Faraday) that leverages coding agents as tools. The judge's low noise and agreement with human assessment suggest reliable automated evaluation for replication quality. Faraday's outperformance of Claude Opus 4.8 and GPT-5.5 on held-out tasks indicates effective generalization beyond training distribution.
This work signals progress in automating scientific replication, a labor-intensive and error-prone process. If such agents become reliable, they could accelerate verification of published research and reduce the replication crisis burden. The use of a 27B-parameter model suggests that specialized post-training on structured tasks can yield competitive performance without the largest frontier models, potentially lowering cost barriers for scientific AI tools.
Potential business value lies in offering automated paper replication services to publishers, universities, and R&D labs, reducing time and cost of verifying results. The rubric-based judge could be commercialized as a quality assessment tool. The approach may also enable faster validation of AI-generated research claims, supporting trust in scientific literature.
Next observable signals include: independent replication of Faraday's results on other benchmarks; release of the Replica task space or judge; adoption by research institutions or journals for automated replication checks; and extension to open-ended scientific innovation tasks. Watch for follow-up papers scaling the approach or applying it to other scientific domains.