Official conference reviewer guidelines outperform LLM-generated reviewer-imitating guidelines in automated peer review study
A new arXiv preprint finds that official conference review criteria yield automated reviews more consistent with human judgments than LLM-generated reviewer-imitating guidelines, and strict rubric-style scoring degrades performance.
1 source · cross-referenced
- Official conference reviewer guidelines produce automated peer reviews that align more closely with human judgments than LLM-generated reviewer-imitating guidelines.
- Strict rubric-style scoring consistently degraded automated review performance, suggesting subjective and holistic scoring is preferable.
- The study evaluates how different reviewer guideline designs affect LLM-based automated peer review systems.
A new arXiv preprint examines how different reviewer guideline designs influence the performance of LLM-based automated peer review systems. The study compares official conference reviewer guidelines with LLM-generated reviewer-imitating guidelines derived from high-quality human reviews. According to the authors, official conference guidelines produced review results that were more consistent with human judgments than the LLM-generated alternatives.
The research further finds that enforcing strict rubric-style scoring consistently degraded performance in automated reviews. This suggests that allowing subjective and holistic scoring may be more effective for LLM-based peer review systems. The authors propose that evaluation criteria refined through established conference practices serve as better guidance for automated reviewing than guidelines designed to imitate human reviewers.
The preprint, titled 'Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review,' was submitted to arXiv on May 16, 2026, and includes 18 pages with two figures. The work is positioned within the ACL 2026 Findings track, indicating it is intended as a contribution to the upcoming Association for Computational Linguistics conference.
- Jul 28, 2026 · arXiv cs.CL
Researchers release GAND, a benchmark to study gender bias in machine translation through gender-ambiguous natural data
Trust79 - Jul 28, 2026 · arXiv cs.CL
Researchers release MioFFAn, an open-source framework for annotating and formalizing scientific formulas with LLM-assisted workflows
Trust79 - Jul 28, 2026 · arXiv cs.AI
Researchers propose SeT-Diff, a diffusion-based foundational model for HPC telemetry and time-series
Trust79