AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

📄 arXiv: 2607.25881v1 📥 PDF

作者: Jia Liu, Veena Krishnaraj, Kateryna Vovk, Kosuke Aizawa, Adrian E. Bayer, Linda Blot, Jessica Cowell, Suyog Garg, Jonathan Grée, Anamaria Hell, Ben Horowitz, Masaya Ichikawa, Kanyuni Iemoto, Keigo Kondo, Zacharie Lorsin, Kevin McCarthy, Jamie Robinson, Miguel Ruiz-Granda, Leander Thiele, Ievgen Vovk, Mingshen Zhou

分类: cs.CL, astro-ph.CO, astro-ph.IM, cs.HC, gr-qc

发布日期: 2026-07-28

备注: 16 pages, 4 figures


💡 一句话要点

探讨大型语言模型在科学项目规划与评估中的应用

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 科学项目规划 提案评估 人工智能 盲评

📋 核心要点

  1. 现有方法在科学项目规划和评估中缺乏有效的工具,尤其是在处理复杂的科学主题时。
  2. 本研究通过比较人类和LLMs生成的项目提案,探讨LLMs在科学研究中的辅助能力。
  3. 实验结果表明,AI评审者对AI生成的提案评分普遍高于人类评审者,且AI评审者的分类准确率达100%。

📝 摘要(中文)

本研究调查了大型语言模型(LLMs)在科学项目规划和提案评估中的有效性。研究中,八个物理、天体物理和宇宙学领域的专家构思的研究项目生成了一页项目计划,分别由人类研究者和三种现代LLMs(ChatGPT、Claude和DeepSeek)独立生成。32个提案由四位人类评审和两种新型LLMs(Claude Opus 4.8和ChatGPT Pro 5.5)进行盲评。结果显示,当前的LLMs能够生成与人类撰写的项目计划相当的提案,但AI评审者对AI生成的提案表现出系统性的偏好。这一发现提示在提案准备和评估中广泛使用LLMs时需谨慎。

🔬 方法详解

问题定义:本研究旨在解决大型语言模型在科学项目规划和提案评估中的有效性问题。现有方法在处理复杂科学主题时,缺乏高效的评估工具,导致评估结果的主观性和不一致性。

核心思路:研究通过生成和评估人类与AI生成的项目提案,探索LLMs在科学研究中的应用潜力。设计上,采用多种LLMs生成提案,并通过人类和AI评审者进行盲评,以确保评估的客观性。

技术框架:整体流程包括项目提案的生成、盲评和结果分析。主要模块包括提案生成(人类与LLMs)、评审(人类与AI)和结果统计与分析。

关键创新:本研究的创新点在于系统性地比较人类与AI生成的提案,并分析不同评审者的评分偏好。这种方法为评估LLMs在科学研究中的应用提供了新的视角。

关键设计:在实验中,使用了四个评估维度的评分标准,确保评审的全面性。同时,采用了最新的LLMs版本,以保证生成提案的质量和相关性。评审者的评分采用五分制,确保了结果的可比性。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,人类评审者对人类和AI生成的提案评分相似,而AI评审者则对AI生成的提案评分高出约一分(满分五分)。人类评审者对提案的识别准确率为72%和79%,而AI评审者的分类准确率达100%。这些结果表明LLMs在项目规划中的有效性和潜在偏见。

🎯 应用场景

该研究的潜在应用领域包括科学研究的项目规划、提案撰写和评估等。通过有效利用LLMs,研究人员可以提高项目提案的质量和效率,进而推动科学研究的进展。未来,LLMs可能在更广泛的科学领域中发挥重要作用,提升研究的自动化和智能化水平。

📄 摘要(原文)

We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.