Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement
作者: Patrick Vossler, Jialin Ouyang, F. Richard Guo, Anran Huang, Ali Shojaie, Lucas Zier, Fan Xia, Jean Feng
分类: stat.ME, cs.AI, stat.AP
发布日期: 2026-08-11
💡 一句话要点
提出专家引导的g-计算方法以优化医院干预效果评估
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 因果推断 医院质量改进 专家引导 g-计算 甘特图 医疗流程优化 大语言模型
📋 核心要点
- 现有方法在评估医院干预措施的因果效应时,面临数据不足和复杂因果机制的挑战。
- 提出的专家引导g-计算方法结合了专家判断与数据驱动模型的优点,提升了因果效应的估计准确性。
- 在模拟实验中,该方法优于传统因果推断方法,并在实际应用中与专家结果高度一致。
📝 摘要(中文)
医院质量改进(QI)项目面临多种候选干预措施以优化医院流程,但现有方法在估计和排名这些干预的因果效应时存在困难。本文聚焦于医院常用指标——平均住院时间(LOS)及其因果估计。我们提出专家引导的g-计算方法,结合了定性和定量方法的优点,通过将甘特图与因果DAG文献相连接,建立了一个因果模型。该方法在模拟中表现优于传统因果推断方法,并在城市安全网医院的研究中与人类专家的结果高度一致。
🔬 方法详解
问题定义:本文旨在解决医院质量改进中干预措施因果效应估计的不足,现有方法在面对假设性干预或复杂因果机制时表现不佳。
核心思路:提出的专家引导g-计算方法通过结合专家输入与数据驱动模型,克服了定性和定量方法的局限性,旨在提高因果效应的估计准确性。
技术框架:该方法包括建立因果模型、专家输入的整合以及基于甘特图的因果推断流程,形成一个系统的估计框架。
关键创新:最大创新在于引入专家引导的g-计算方法,通过仅在数据无法识别的部分寻求专家输入,显著提高了因果效应的估计能力。
关键设计:在技术细节上,设计了一个LLM辅助的管道,以可靠地扩展专家推理,确保在多样化的因果结构和干预机制下的有效性。
🖼️ 关键图片
📊 实验亮点
在模拟实验中,专家引导的g-计算方法在患者具有多样化因果结构和干预机制时,表现出比传统因果推断方法更优的性能,且在城市安全网医院的实际研究中,生成的图表和时间节省估计与人类专家结果高度一致。
🎯 应用场景
该研究的潜在应用领域包括医院质量改进、医疗流程优化等,能够为医院管理者提供科学的决策支持,提升医疗服务质量。未来,该方法也可扩展到其他领域的因果效应估计,具有广泛的实际价值。
📄 摘要(原文)
Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal effects of such interventions. This work focuses on one of the most standard hospital metrics, the average length of stay (LOS), and its causal estimand, the average time saved. To characterize this causal effect, qualitative approaches rely on expert judgment to map patient trajectories, making them susceptible to cognitive biases; quantitative approaches rely on data-driven models, which fail when interventions are hypothetical with no historical data or have complex causal mechanisms that require clinical reasoning rather than data alone. We propose expert-guided g-computation, or egg-computation, which combines the complementary strengths of both approaches by connecting the Gantt charts commonly used to map patient trajectories with the causal DAG literature. We introduce a causal model over Gantt charts and establish identification using a variant of g-computation that seeks expert input only for components unidentifiable from data. To make egg-computation practical, we develop an LLM-assisted pipeline that reliably scales up expert reasoning. In simulations, egg-computation outperforms conventional causal inference methods when patients have diverse causal structures and intervention mechanisms. In a study of eleven candidate QI interventions at an urban safety-net hospital, the LLM pipeline generated graphs and time-saving estimates highly concordant with those of human experts. Beyond healthcare, egg-computation is a broadly applicable framework for estimating the average time saved for candidate interventions whose causal mechanisms can be represented using Gantt charts.