How Closely Do LLM Reviews Align with Human Peer Review?
A study compared reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 ICLR 2026 submissions (100 oral, 100 poster, 100 rejected). All LLMs distinguished accepted from rejected papers but failed to reproduce the oral vs. poster distinction. Gemini assigned systematically higher ratings; OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers.
A cross-provider analysis evaluated how closely LLM-generated reviews align with human peer review using 300 ICLR 2026 submissions. Three models—OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6—reviewed each paper under identical instructions. All models could separate accepted from rejected papers, but none captured the finer oral/poster distinction. Rating behaviors varied: Gemini was more generous, while OpenAI and Claude were harsher on oral papers but aligned better with humans on rejected and poster papers.
The study reveals that current LLMs can approximate binary accept/reject decisions but lack the granularity to replicate nuanced human distinctions like oral vs. poster. Provider-specific calibration differences suggest that raw LLM scores are not directly interchangeable with human ratings without adjustment.
As LLMs are increasingly used to assist or automate scientific peer review, this evidence highlights both promise and limitations. While they can serve as a first-pass filter, their systematic biases and inability to capture fine-grained quality tiers mean they are not yet ready to replace human judgment in high-stakes conference settings.
For publishers and conference organizers, LLM-based review tools could reduce reviewer workload by triaging submissions, but careful calibration and human oversight remain essential to avoid misclassification of borderline papers. Providers may differentiate by offering better-aligned review APIs.
Future work may focus on calibrating LLM reviewers to match human distributions, exploring ensemble approaches across providers, and investigating whether fine-tuning on domain-specific review data can improve alignment with nuanced decision categories.