Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

📄 arXiv: 2608.26885v1 📥 PDF

作者: Nikol Figalová, Lynn Huestegge, Anne Böckler-Raettig

分类: cs.AI, cs.HC, cs.SE

发布日期: 2026-08-27


💡 一句话要点

比较人类与LLM筛选工作流程以优化文献综述中的回忆与工作负载平衡

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 文献筛选 系统评价 回忆率 工作负载 人类监督 机器学习

📋 核心要点

  1. 现有的文献筛选方法在处理复杂概念时,常常面临假阴性问题,导致相关研究被遗漏。
  2. 本研究通过比较人类与LLM的筛选工作流程,提出了优化回忆与工作负载平衡的解决方案。
  3. 实验结果表明,虽然没有工作流程能完全恢复所有合格记录,但LLM在高回忆任务中表现出色,尤其是在经过验证的人类监督工作流程中。

📝 摘要(中文)

背景:大型语言模型(LLMs)在证据合成中的筛选应用日益增多,但假阴性可能导致相关研究在全文评估前被遗漏。我们在一项预注册研究中比较了人类与LLM的标题和摘要筛选工作流程。方法:在保守的标题筛选后,1,131条记录由一名审查负责人、四名训练助手和七次不同模型及处理配置的LLM运行进行筛选。结果显示,无一工作流程能恢复所有已验证的合格记录。人类工作流程和两次GPT-5.4文件批次运行保留了42.2-45.0%的记录,回忆率为82.3-82.9%。讨论指出,LLM的筛选性能依赖于实施的工作流程,而非单一模型身份。

🔬 方法详解

问题定义:本研究旨在解决在复杂概念文献综述中,现有筛选方法导致的假阴性问题,尤其是如何在回忆率与工作负载之间找到平衡。

核心思路:通过比较人类与不同配置的LLM筛选工作流程,探索在高回忆任务中如何优化筛选效率和准确性。

技术框架:研究分为几个阶段:首先进行保守的标题筛选,然后由审查负责人和训练助手进行非重叠子集的筛选,最后使用不同模型和配置的LLM进行筛选。

关键创新:本研究的创新点在于强调了工作流程的设计对LLM筛选性能的影响,指出仅依赖模型身份不足以提升筛选效果。

关键设计:实验中使用了不同的处理配置,包括文件批次和一次性处理,比较了不同模型(如GPT-5.4和Gemini 3.1)的表现,重点关注回忆率和记录保留率。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,所有工作流程均未能恢复所有已验证的合格记录。人类工作流程和两次GPT-5.4文件批次运行保留了42.2-45.0%的记录,回忆率为82.3-82.9%。Gemini 3.1文件批次的回忆率最高,达83.9%,但保留的记录比例为56.7%。

🎯 应用场景

该研究的结果对文献综述、系统评价及其他需要筛选大量文献的领域具有重要应用价值。通过优化人类与LLM的协作工作流程,可以提高筛选效率,减少遗漏相关研究的风险,进而提升科学研究的质量和效率。

📄 摘要(原文)

Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review. Methods. After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational. Results. No workflow recovered all verified eligible records. The human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records while achieving 82.3-82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All-at-once configurations recovered fewer eligible records than corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run. Discussion. LLM screening performance depended on the implemented workflow, not model identity alone. Processing configuration, workload, record-level variation, and human-LLM decision integration are therefore substantive properties of deployed systems. For high-recall tasks, LLMs are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.