What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation
作者: Ziyue Wang, Aomufei Yuan, Yiran Yao, Linli Yao, Hongyao Zuo, Ziwen Gong, Yuanxin Liu, Shicheng Li, Yishuo Cai, Tong Yang, Xu Sun, Xiaohui Li, Haoli Bai
分类: cs.CL, cs.AI
发布日期: 2026-08-24
备注: Equal contribution by Ziyue Wang, Aomufei Yuan and Yiran Yao. Corresponding authors: Tong Yang and Xu Sun
💡 一句话要点
提出Lit2Test基准以评估语言模型的研究创意质量
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 语言模型 研究创意 可证伪性 评估基准 科学发现 模型比较 客观评估
📋 核心要点
- 现有语言模型在提出研究创意时缺乏统一的评判标准,导致评估结果不一致。
- 本文提出Lit2Test基准,通过预设可证伪的结果,使研究提案的质量可判定。
- 实验结果显示,Lit2Test能够在10000次重采样中严格排名四个模型,且结果基于测试质量而非表面流畅性。
📝 摘要(中文)
大型语言模型在提出研究创意方面的应用日益增多,但现有评判方法缺乏统一的决策规则。本文提出Lit2Test基准,通过六个领域的合同围绕可证伪的结果组织,使每个提案预先承诺能够证明其错误的观察,从而使其质量可判定而非仅仅可争论。Lit2Test基于200个真实论文邻域构建,收集四个前沿模型的提案,并通过1200次盲评进行比较。该基准审计自身的可靠性,并在所有10000次自助重采样中恢复出四个模型的严格排名,分离来自于提出的测试和指标的质量,而非表面流畅性。我们公开发布该基准及其构建流程和审计文档。
🔬 方法详解
问题定义:本文旨在解决现有语言模型在研究创意评估中的主观性和不一致性问题,现有方法往往依赖于风格和立场,缺乏客观标准。
核心思路:Lit2Test基准通过设定一个围绕可证伪结果的六个领域合同,使得每个研究提案都能预先承诺能够证明其错误的观察,从而使质量评估变得可判定。
技术框架:该基准从200个真实论文邻域构建,收集四个前沿模型的提案,并进行1200次盲评比较。整个流程包括提案收集、盲评、结果分析和可靠性审计等模块。
关键创新:Lit2Test的最大创新在于其通过可证伪性来定义研究提案的质量,使得评估不再仅仅依赖于主观判断,而是基于明确的可验证结果。
关键设计:在设计中,采用了明确的评估标准和可靠性审计机制,确保评估过程的客观性和一致性,且通过三位注释者的协作来验证结果的可靠性。
🖼️ 关键图片
📊 实验亮点
实验结果表明,Lit2Test在所有10000次自助重采样中成功恢复了四个模型的严格排名,且这种排名的分离主要源于提出的测试和指标的质量,而非表面流畅性,显示出该基准的有效性和可靠性。
🎯 应用场景
该研究的潜在应用领域包括学术研究、科学发现和创新提案等。Lit2Test基准能够为研究人员提供一个客观的评估工具,帮助他们更有效地提出和验证研究创意,从而推动科学研究的进展。
📄 摘要(原文)
Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.