How Closely Do LLM Reviews Align with Human Peer Review?
作者: Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez
分类: cs.CL, cs.AI
发布日期: 2026-08-04
💡 一句话要点
比较大型语言模型与人类评审的一致性
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 科学评审 人类评审 决策一致性 评审优先级 ICLR 2026
📋 核心要点
- 现有评估方法未能充分考察不同大型语言模型与人类评审之间的一致性,尤其是在细致决策层面。
- 本文通过比较三种大型语言模型的评审结果与人类评审,探讨其在决策一致性和评审优先级上的差异。
- 实验结果显示,尽管LLM能够区分接受与拒绝的论文,但在口头与海报的区分上表现不佳,且强调点存在显著差异。
📝 摘要(中文)
大型语言模型(LLMs)在生成科学评审方面的应用日益增加,但现有评估很少在相同的控制环境中考察不同模型与会议决策及人类评审优先级的一致性。本文比较了OpenAI GPT-5.4、Google Gemini 3.1 Pro Preview和Anthropic Claude Opus 4.6与人类评审及300篇主题匹配的ICLR 2026提交论文的最终决策。研究贡献在于对三种模型在广泛和细致决策类别的一致性、推荐尺度使用差异及识别弱点的主题一致性进行了交叉分析。结果显示,所有三种LLM能够区分接受与拒绝的论文,但未能重现人类评审中的口头与海报区分。LLM与人类评审在强调点上存在差异,LLM更常识别缺失的基线比较,而人类则更关注计算效率问题。
🔬 方法详解
问题定义:本文旨在解决大型语言模型在科学评审中与人类评审一致性不足的问题,尤其是在细致的决策层面。现有方法未能有效评估不同模型的评审结果与人类评审之间的关系。
核心思路:通过对三种大型语言模型的评审结果与人类评审进行系统比较,分析其在决策一致性、推荐尺度使用及识别弱点方面的差异,以揭示LLM在科学评审中的潜在局限性。
技术框架:研究设计包括对300篇ICLR 2026提交论文的评审,论文被分为口头、海报和拒绝三类。每个模型在去除决策信息后,使用相同的指令和评分标准对每篇论文进行评审。
关键创新:本研究的创新在于首次进行跨模型的评审一致性分析,揭示了LLM与人类评审在广泛决策一致性与细致判断上的差异,强调了LLM在科学评审中的局限性。
关键设计:研究中使用的评分尺度和评审指令保持一致,确保了评审结果的可比性。模型的评分模式表现出提供者特异性,Gemini模型评分普遍较高,而OpenAI和Claude在拒绝和海报论文上更接近人类评审。
🖼️ 关键图片
📊 实验亮点
实验结果表明,所有三种LLM能够有效区分接受与拒绝的论文,但未能重现人类评审中的口头与海报区分。Gemini模型的评分普遍高于其他模型,而OpenAI和Claude在拒绝和海报论文上的评分更接近人类评审,显示出不同模型在评审中的特异性。
🎯 应用场景
该研究的潜在应用领域包括科学论文评审、学术会议组织及大型语言模型的优化。通过深入理解LLM与人类评审之间的差异,可以为未来的评审系统设计提供指导,提升评审质量与效率。
📄 摘要(原文)
Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.