Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories

📄 arXiv: 2607.28567v1 📥 PDF

作者: Mengfei Ran, Yifeng Shen, Ruijie Guan

分类: stat.ML, cs.LG, stat.AP, stat.CO, stat.ME

发布日期: 2026-07-30


💡 一句话要点

提出双重稳健功能表示学习以解决不规则历史的因果推断问题

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 因果推断 纵向研究 功能表示学习 不规则历史 双重稳健估计

📋 核心要点

  1. 现有的双重稳健估计器通常依赖标量摘要,无法有效处理不规则历史数据,限制了因果推断的准确性。
  2. 本文提出DR-FRL,通过功能和时间编码器将不规则历史转化为状态,结合干扰头估计多种函数,提升因果推断能力。
  3. 实验结果表明,在高维功能混淆和信息丰富的测量条件下,DR-FRL显著提高了因果推断的准确性,尤其在重尾伪结果情况下表现优异。

📝 摘要(中文)

纵向因果研究通常记录不规则的功能片段历史,如实验室值、生理信号、传感器流和不等时点的图像摘要。标准的双重稳健估计器通常需要标量摘要,而序列学习者优化的预测损失未必能稳定有效影响函数。本文提出双重稳健功能表示学习(DR-FRL),通过交叉拟合工作流将不规则历史转化为针对估计量的状态。功能和时间编码器将点云和先前历史映射到状态;干扰头估计结果、处理和审查函数;并通过有效影响函数(EIF)目标验证、校准、重叠、尾部和消融诊断评估状态是否支持估计方程。模拟结果表明,在功能混淆高维、测量信息丰富、支持弱或伪结果重尾的情况下,DR-FRL显示出显著的提升。

🔬 方法详解

问题定义:本文旨在解决在不规则历史数据下进行纵向因果推断的挑战。现有方法通常依赖标量摘要,无法充分利用丰富的历史信息,导致因果推断的准确性降低。

核心思路:论文提出的DR-FRL通过将不规则历史转化为针对估计量的状态,利用功能和时间编码器来映射历史数据,结合干扰头来估计结果、处理和审查函数,从而提升因果推断的稳健性和准确性。

技术框架:DR-FRL的整体架构包括多个模块:功能和时间编码器负责将点云和历史数据映射为状态;干扰头用于估计结果、处理和审查函数;最后,通过EIF目标进行验证和诊断,确保状态支持估计方程。

关键创新:最重要的技术创新在于将不规则历史数据有效转化为状态,并通过EIF目标进行校准和验证,确保了估计的稳健性。这一方法与传统依赖标量摘要的估计器本质上不同,能够更好地处理复杂的历史数据。

关键设计:在设计中,采用了特定的损失函数来优化状态的表示,同时设置了多个干扰头以估计不同的函数。网络结构上,功能和时间编码器的设计考虑了数据的时序特性,以提高模型的表达能力。实验中还采用了Catoni聚合作为有界影响点估计器,确保了估计的稳定性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,在高维功能混淆和信息丰富的测量条件下,DR-FRL的因果推断准确性显著提升,尤其在伪结果重尾情况下,表现出比传统方法更优的性能,验证了其有效性。

🎯 应用场景

该研究的潜在应用领域包括医疗健康、临床试验和社会科学等领域,尤其是在处理不规则时间序列数据时,能够提供更为准确的因果推断。这将有助于改善临床决策和政策制定,推动个性化医疗的发展。

📄 摘要(原文)

Longitudinal causal studies often record histories as irregular functional fragments: laboratory values, physiologic signals, sensor streams, and image-derived summaries measured at unequal and informative times. Standard doubly robust estimators usually require scalar summaries, whereas sequence learners optimize prediction losses that need not stabilize the efficient influence function. We propose Doubly Robust Functional Representation Learning (DR-FRL), a cross-fitted workflow that turns irregular histories into estimand-targeted states for observed-history regimes. Functional and temporal encoders map point clouds and prior histories into states; nuisance heads estimate outcome, treatment, and censoring functions; and EIF-targeted validation, calibration, overlap, tail, and ablation diagnostics assess whether the state supports the estimating equation. If the selected state preserves the nuisance information needed by the EIF, representation error enters the same second-order product remainder as ordinary nuisance error, and the mean estimator is asymptotically linear under explicit rate, overlap, calibration, and stability conditions. Catoni aggregation is treated separately as a bounded-influence point estimator, not a replacement for Wald inference. Simulations show gains when functional confounding is high-dimensional, measurement is informative, support is weak, or pseudo-outcomes are heavy-tailed. A VitalDB audit shows that DR-FRL can use irregular laboratory point clouds and deliver a useful negative finding: for this ICU-disposition endpoint, scalar laboratory summaries already carry much endpoint-relevant information.