How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

📄 arXiv: 2608.08975v1 📥 PDF

作者: Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou

分类: cs.CL, cs.AI

发布日期: 2026-08-10


💡 一句话要点

探讨修辞选择如何影响AI评审判断以应对奖励黑客问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: AI评审 修辞选择 奖励黑客 科学评估 大型语言模型 评审机制 学术出版

📋 核心要点

  1. 现有的AI评审系统在处理科学内容时,可能受到修辞选择的影响,导致评审结果的不一致性。
  2. 本文提出通过构建受控语料库和使用LLM重写器,系统性地研究修辞选择对AI评审判断的影响。
  3. 实验结果显示,修辞敏感性存在结构化特征,且不同的修辞维度对评审结果的影响程度各异。

📝 摘要(中文)

随着大型语言模型在科学评估中的参与日益增加,本文研究了一种潜在的奖励黑客形式:修辞选择如何在科学内容保持不变的情况下影响AI评审判断,以及这些影响在不同评估条件下的变化。我们构建了一个包含4200篇完整论文的受控语料库,并通过两种LLM重写器对六个修辞维度进行对立方向的转换,随后由五个LLM评审在标准和严格协议下评估结果。研究结果表明,修辞敏感性是结构化的,证据框架和新颖性立场在整体评估中产生了最大的正负对比,而其他维度的影响较小或不稳定。这些发现为评估系统提供了对科学写作中内容保持变化的鲁棒性。

🔬 方法详解

问题定义:本文旨在解决修辞选择对AI评审判断的影响问题,现有方法在评审过程中未能充分考虑修辞因素的作用,导致评审结果的不一致性和潜在的奖励黑客现象。

核心思路:通过构建一个包含4200篇论文的受控语料库,利用两种LLM重写器对六个修辞维度进行对立方向的转换,从而系统性地分析修辞选择对AI评审的影响。

技术框架:整体流程包括数据收集、修辞重写、AI评审以及结果分析。首先,构建语料库,然后通过重写器对论文进行修辞转换,最后由LLM评审进行评估。

关键创新:本文的主要创新在于系统性地揭示了修辞敏感性是结构化的,而非均匀的,且不同修辞维度对评审结果的影响程度存在显著差异。

关键设计:在实验中,采用了标准和严格的评审协议,重写器的选择和修辞维度的设置是关键参数,评审结果的分析基于AI评审者的原始评分,确保了结果的可靠性和有效性。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果表明,在严格评审条件下,平均OA评分下降了1.36分,但修辞敏感性未发生显著变化。证据框架和新颖性立场在评审中产生了最大的正负对比,显示出修辞选择对AI评审的显著影响。

🎯 应用场景

该研究的潜在应用领域包括科学论文评审、学术出版以及AI辅助的评估系统。通过理解修辞选择对评审结果的影响,可以设计出更为公正和透明的评审机制,从而提升学术交流的质量和效率。

📄 摘要(原文)

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.