MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

📄 arXiv: 2607.25186v1 📥 PDF

作者: Xiao Li, Mouxiao Bian, Zhaodi Wu, Sijie Ren, Juechen Chen, Lu Lu, Jingru Ding, Yun Zhong, Jie Xu, Yixiu Liang, Junbo Ge

分类: cs.CL

发布日期: 2026-07-28


💡 一句话要点

提出MyoCardBench以解决心血管护理场景下LLM评估问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 心血管护理 真实世界基准 多任务评估 临床决策支持

📋 核心要点

  1. 现有医学LLM基准测试多集中于孤立任务,缺乏对心血管护理复杂流程的全面评估。
  2. MyoCardBench通过整合多种心血管护理任务,提供了一个真实世界的评估框架,涵盖了临床维度。
  3. 实验结果显示,GPT-5.4在多项评估中表现优异,宏平均得分达到62.55,显著高于其他模型。

📝 摘要(中文)

背景:大多数医学大型语言模型(LLM)基准测试集中于考试知识或孤立任务,无法反映心血管护理的纵向、多模态和安全关键工作流程。目标:开发MyoCardBench,一个涵盖心血管护理连续体的真实世界基准,并评估LLM在临床维度和专业任务中的表现。方法:MyoCardBench包含来自13个任务特定数据集的2263个项目,数据源于去标识化的心血管记录和检查数据。16名心脏病医生进行了注释和参考构建,随后由两名资深心脏病专家进行交叉审查。七个LLM在标准化的零样本设置下生成15841个输出。结果:GPT-5.4在所有三个维度中排名第一,宏平均得分为62.55。结论:MyoCardBench是迄今为止最大的真实世界多任务基准,提供了对临床真实心脏病场景的广泛覆盖。

🔬 方法详解

问题定义:论文要解决的问题是现有医学LLM基准测试无法全面反映心血管护理的复杂性和多样性,导致模型评估不够准确。

核心思路:论文提出MyoCardBench,通过整合来自真实世界的心血管护理数据,构建一个多任务基准,旨在全面评估LLM在临床场景中的表现。

技术框架:MyoCardBench的整体架构包括数据收集、注释、模型生成和评估四个主要阶段。数据来源于去标识化的心血管记录,经过专业医生的注释和审核后,供LLM生成输出。

关键创新:MyoCardBench的最大创新在于其覆盖了心血管护理的整个连续体,提供了多任务的真实世界基准,填补了现有基准的空白。

关键设计:在数据集构建中,采用了2263个项目,涵盖13个任务特定数据集,评估指标包括关键点覆盖率和整体临床质量,确保评估的全面性和准确性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,GPT-5.4在MyoCardBench中表现最佳,宏平均得分为62.55,领先于Gemini 3.1 Pro(59.95)和Qwen 3.6 27B(59.72)。在具体任务中,CardioAuxReport得分最高(86.38),而CardioECGRead和CardioEthics得分最低,显示出不同任务间的显著差异。

🎯 应用场景

MyoCardBench的研究成果可广泛应用于心血管医学领域,帮助医生和研究人员评估和优化大型语言模型在临床决策支持中的应用潜力。未来,该基准还可能推动LLM在其他医学领域的应用,提升医疗服务的质量和效率。

📄 摘要(原文)

Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop MyoCardBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: MyoCardBench includes 2,263 items from 13 task-specific datasets derived from de-identified cardiovascular records and examination data. Sixteen cardiology physicians conducted annotation and reference construction, followed by cross-review from two senior cardiologists. Seven LLMs generated 15,841 outputs under standardized zero-shot settings. Open-ended tasks were evaluated using key-point coverage and holistic clinical quality, while CardioEthics was scored by accuracy. Results: GPT-5.4 achieved the highest macro-average (62.55) and item-weighted mean (62.19), followed by Gemini 3.1 Pro (59.95) and Qwen 3.6 27B (59.72). GPT-5.4 ranked first in all three dimensions. CardioAuxReport performed best (86.38), whereas CardioECGRead (17.25) and CardioEthics (17.34) were lowest. The largest gaps between holistic clinical quality and key-point coverage occurred in CardioComm (52.71), CardioEmergRescue (52.05), and CardioTreatPlan (48.80). Conclusions: To our knowledge, MyoCardBench is the largest real-world, multi-task benchmark for LLM evaluation across the cardiovascular care continuum and offers the broadest coverage of clinically authentic cardiology scenarios reported to date. It provides a rigorous framework for identifying model strengths, clinically important omissions, and priorities for future development.