AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

📄 arXiv: 2607.25672v1 📥 PDF

作者: Anamaria Hell, Kateryna Vovk, Veena Krishnaraj, Jia Liu, Kosuke Aizawa, Adrian E. Bayer, Linda Blot, Jessica Cowell, Suyog Garg, Jonathan Grée, Ben Horowitz, Masaya Ichikawa, Kanyuni Iemoto, Keigo Kondo, Zacharie Lorsin, Kevin McCarthy, Jamie Robinson, Miguel Ruiz-Granda, Leander Thiele, Ievgen Vovk, Mingshen Zhou

分类: astro-ph.IM, astro-ph.CO, cs.CL, gr-qc

发布日期: 2026-07-28

备注: 12 pages, 3 figures


💡 一句话要点

评估大型语言模型在科学文献综述中的辅助能力

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 文献综述 科学研究 物理学 天体物理学 宇宙学 AI辅助 文献选择

📋 核心要点

  1. 现有大型语言模型在科学文献综述中表现不足,无法有效替代人类专家的文献搜索能力。
  2. 通过对比人类与AI在文献选择上的差异,评估AI在文献综述中的辅助作用和可靠性。
  3. 研究发现2026年模型在性能上有显著提升,特别是在减少虚构和元数据错误方面。

📝 摘要(中文)

本研究探讨了大型语言模型(LLMs)在物理学、天体物理学和宇宙学领域文献综述中的辅助能力。通过对八个专家设计的研究项目进行对比实验,发现人类专家与AI模型所选文献的重叠率低于6%,表明AI尚未能独立进行有效的文献搜索。研究还评估了AI生成的参考文献的可靠性,发现3%的文献为虚构,64%的文献存在元数据错误。2026年的ChatGPT Pro 5.5模型在单项目测试中表现显著提升,无虚构或元数据错误。

🔬 方法详解

问题定义:本研究旨在解决大型语言模型在科学文献综述中的有效性问题,现有模型在文献选择上与人类专家存在显著差距。

核心思路:通过对八个研究项目进行对比实验,评估AI与人类在文献选择上的一致性,并分析AI生成文献的可靠性。

技术框架:研究设计包括文献选择任务的控制实验,涉及人类专家和不同版本的AI模型(如ChatGPT-4o、ChatGPT Deep Research和Gemini),并对结果进行系统比较。

关键创新:本研究首次系统性地评估了大型语言模型在科学文献综述中的应用潜力,揭示了AI生成文献的虚构和元数据错误问题。

关键设计:研究中对AI生成的文献进行了分类,识别出3%的虚构文献和64%的元数据错误,强调了对AI生成内容进行系统验证的必要性。具体参数和模型版本的选择也对结果产生了影响。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,AI与人类专家在文献选择上的重叠率低于6%,而2026年模型ChatGPT Pro 5.5在单项目测试中表现出色,未出现虚构或元数据错误,显示出显著的性能提升。

🎯 应用场景

该研究的结果为科学研究中的文献综述提供了新的视角,表明大型语言模型可以作为人类专家的辅助工具,尤其在文献筛选和初步分析阶段。未来,随着模型的不断改进,AI在科学研究中的应用潜力将进一步扩大。

📄 摘要(原文)

We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.