Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
作者: Zheng Tong, Yang Liu, Wanshu Fan, Jing Qin, Zhongbin Han, Haifan Gong, Congyu Liao, Xiaofeng Liu, Cong Wang
分类: cs.CV
发布日期: 2026-07-28
备注: Review article, 6 figures, 2 tables. 42 pages
💡 一句话要点
提出代理智能以解决医学领域多步骤临床任务的挑战
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 代理智能 医学人工智能 多步骤任务 临床转化 评估框架 文献回顾 系统设计
📋 核心要点
- 现有医学人工智能系统在执行复杂临床任务时面临规划、工具使用和多代理协调等挑战。
- 论文通过系统文献回顾,提出了代理智能的框架,旨在提升医学AI在多步骤任务中的应用能力。
- 研究纳入557项相关研究,表明当前的评估方法与临床需求不匹配,强调了未来研究的方向。
📝 摘要(中文)
大型语言模型和多模态基础模型使医学人工智能系统能够超越孤立预测,执行需要规划、工具使用、记忆、迭代修正和专业代理协调的多步骤临床任务。然而,医学领域代理智能的范围尚未确定,当前的评估实践与临床使用要求尚不一致。我们进行了系统的文献回顾,筛选了1,649条记录,最终纳入557项符合预定义标准的独特研究。这些研究描述了使用外部工具的单一代理、支持检索和外部知识的工作流程、多模态代理及多代理系统在医学问答、图像解读、电子健康记录分析、药物安全和临床试验预测中的应用。现有证据主要集中在公共基准、模拟环境、回顾性数据集和小规模专家评估上,过程可靠性、证据可追溯性、不确定性、安全性、工作流程影响和外部有效性评估不够一致。临床转化将依赖于更清晰的定义、可重复的评估、可审计的监督、可互操作的系统设计和在真实临床工作流程中的前瞻性验证。
🔬 方法详解
问题定义:本论文旨在解决医学领域中代理智能的应用范围不明确和评估实践与临床需求不一致的问题。现有方法在多步骤临床任务中缺乏有效的规划和协调能力。
核心思路:论文提出通过系统文献回顾和证据映射,明确代理智能在医学中的应用场景和评估标准,以促进其临床转化。
技术框架:整体架构包括文献筛选、证据分类和评估标准制定三个主要模块。文献筛选阶段筛选出符合条件的研究,证据分类阶段对不同类型的代理智能应用进行归类,评估标准制定阶段则提出可操作的临床评估标准。
关键创新:最重要的创新点在于系统性地整合了多种代理智能的应用案例,并提出了针对临床转化的评估框架,与现有方法相比,更加注重实际应用中的可行性和有效性。
关键设计:在文献筛选中,采用了预定义标准以确保研究的相关性;在评估标准中,强调了过程可靠性和外部有效性等关键指标,以确保代理智能在临床环境中的适用性。
🖼️ 关键图片
📊 实验亮点
研究纳入557项相关研究,表明现有医学AI系统在多步骤任务中的评估方法不够一致。通过系统文献回顾,提出了新的评估框架,强调了过程可靠性和外部有效性,为未来的临床转化提供了重要参考。
🎯 应用场景
该研究的潜在应用领域包括医学问答、图像解读、电子健康记录分析等,能够显著提升医疗决策的效率和准确性。未来,代理智能有望在临床工作流程中发挥更大作用,推动个性化医疗和智能化管理的发展。
📄 摘要(原文)
Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.