Event date · · arXiv

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

FACT STATEMENT

A study surveyed AI policies at 111 AI/NLP conferences and medical journals, finding substantial regulation differences. It evaluated AI-generated reviews at ICLR 2026 and Nature Communications using a dataset of original submissions and hundreds of human and machine reviews. Current LLMs produce detailed, fluent reviews but show systematic weaknesses: overly positive recommendations, generic criticism, and uneven evidence grounding. Aggregate quality scores alone can overestimate review quality.

What happened

A study published on arXiv on August 4, 2026, examined AI-assisted peer review by surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing significant regulatory differences between the two communities. The research also evaluated AI-generated peer reviews at ICLR 2026 and Nature Communications using a novel dataset of original manuscript submissions and several hundred human- and machine-generated reviews. Comparisons of open-source and proprietary models using metrics such as LLM-as-a-Judge, score alignment, granularity, and overlap with human concerns showed that while current LLMs can generate detailed and fluent reviews, they exhibit systematic weaknesses including overly positive recommendations, generic criticism, and uneven evidence grounding. The study argues that aggregate quality scores alone can overestimate review quality and calls for multi-dimensional evaluation.

Technical significance

LLM-generated reviews achieve high fluency and detail but suffer from poor calibration (overly positive scores), lack of specificity, and insufficient grounding in manuscript evidence. Multi-dimensional evaluation beyond aggregate scores is necessary to assess review quality accurately.

Industry impact

The significant policy divergence between AI/NLP conferences and medical journals suggests a lack of consensus on acceptable AI use in peer review, potentially affecting cross-disciplinary collaboration and trust in AI-assisted review processes.

Decision value

Publishers and conference organizers can use these findings to refine AI policies and develop tools that augment human reviewers while mitigating risks of bias and low-quality reviews. AI developers may target improvements in evidence grounding and score calibration.

What to watch

Future work may focus on developing standardized AI review policies and improving LLM review systems to address identified weaknesses. Observable next signals include updates to conference and journal AI policies, and new benchmarks for AI review quality.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.