Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography
作者: Louis Chen, Torbjörn E. M. Nordling
分类: cs.CV, cs.AI, eess.IV
发布日期: 2026-09-03
备注: 92 pages, 38 figures, incl. supplementary
💡 一句话要点
提出量化解释方法以提升远程光电容积描记法的可靠性
🎯 匹配领域: 支柱八:物理动画 (Physics-based Animation)
关键词: 远程光电容积描记法 模型解释 量化分析 心率估计 数据集迁移 性能评估 深度学习
📋 核心要点
- 现有的远程光电容积描记法解释主要依赖热图,缺乏定量分析,导致解释的可靠性不足。
- 本文提出了一种量化的解释方法,通过评估皮肤覆盖率和SaCo,探索其与模型性能的关系。
- 实验结果表明,Beyond Intuition在不同数据集上的表现优于其他方法,但与心率估计的准确性并无直接关联。
📝 摘要(中文)
背景:远程光电容积描记法通过面部视频估计心血管脉搏,其解释主要依赖热图,而缺乏定量证据。我们量化了解释并探讨其在不同数据集间的迁移性及与模型性能的关系。方法:训练了八个条件特定的RhythmFormer模型,评估了不同光照条件下的心率估计。结果显示,Beyond Intuition在两个数据集上的表现优异,但与心率误差等性能指标并无显著相关性。结论:皮肤覆盖率和SaCo提供了与性能指标互补的信息,归因于皮肤并不保证准确估计。
🔬 方法详解
问题定义:本文旨在解决远程光电容积描记法中解释的定量不足问题,现有方法主要依赖热图,缺乏对模型关注区域的深入分析。
核心思路:通过量化模型的解释,特别是皮肤覆盖率和Salience-guided Faithfulness Coefficient (SaCo),来评估模型在不同条件下的性能表现。
技术框架:研究中训练了八个条件特定的RhythmFormer模型,分别在不同光照、说话、旋转和骑行条件下进行心率估计,比较了不同解释方法的效果。
关键创新:引入了Beyond Intuition作为解释方法,其在两个数据集上表现优异,提供了更可靠的模型关注区域分析,与传统热图方法相比具有更高的解释能力。
关键设计:模型训练过程中,设置了不同的光照条件和运动状态,评估了模型的皮肤覆盖率和SaCo,发现这些指标与心率估计的准确性并无直接相关性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,Beyond Intuition在两个数据集上的中位皮肤覆盖率为0.789和0.826,SaCo分别为0.837和0.917,表现优于其他方法。然而,在40 lux光照条件下,其性能显著下降,显示出环境因素对模型性能的影响。
🎯 应用场景
该研究的潜在应用领域包括远程医疗监测、健康管理和运动生理学等。通过提高远程光电容积描记法的解释性和可靠性,可以更好地为临床决策提供支持,促进个性化医疗的发展。
📄 摘要(原文)
Background. Remote photoplethysmography estimates the cardiovascular pulse from facial video, and its explanations have rested on inspecting heatmaps rather than on quantitative evidence about where a model reads it. We quantified the explanations and asked whether such explanations transfer between datasets and track model performance. Method. We trained eight condition-specific RhythmFormer models on NCKU-rPPG, recorded under three illumination levels, speaking, rotation, and cycling, estimated one heart rate per 5.12-second clip, and set them beside a UBFC-rPPG reproduction. Raw attention, rollout, attention flow, and Beyond Intuition were assessed by skin coverage and the Salience-guided Faithfulness Coefficient (SaCo). Results. Beyond Intuition ranked highest on both datasets, at median coverage 0.789 and SaCo 0.837 on Static level 3 against 0.826 and 0.917 on UBFC-rPPG; lower ranks differed. Within one participant of one condition, neither measure was related to a clip's heart-rate error, waveform correlation, or signal-to-noise ratio on either dataset: 186 of the 252 coefficients fell below $|ρ|=0.10$ and 28 reached $p<0.05$ against the 13 expected by chance. Across the eight scenarios only Beyond Intuition's coverage followed the three performance measures, at $ρ=-0.43$, $+0.57$, and $+0.43$, while the attention-only methods' SaCo ran opposite to each. It failed at 40 lux alone, its median coverage falling to 0.180 and its median SaCo to $-0.178$, whereas motion degraded the estimates far more without such a drop. Conclusions. Skin coverage and SaCo carry information complementary to the performance measures rather than a proxy for them: attributing to the skin does not guarantee an accurate estimate. What an attribution reveals about a condition is where the model looks rather than how faithfully its map is ordered.