检验AI解释是否跨数据集可靠,发现皮肤覆盖不等于准确心率估计
Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography

- 用皮肤覆盖率和萨科系数量化模型注意力位置
- 仅超越直觉方法在多场景下解释与性能相关
- 光照极低时解释失效,但模型仍可工作,说明解释不可靠
远程光电容积脉搏波图(rPPG)通过面部视频估计心率,其解释依赖热力图而非定量证据。本文在NCKU-rPPG数据集上训练八种特定条件的RhythmFormer模型,涵盖三种光照水平、说话、旋转和骑行,每5.12秒输出一次心率,并与UBFC-rPPG复现结果对比。采用原始注意力、回溯法、注意力流及超越直觉方法,通过皮肤覆盖率和显著性引导忠实度系数(SaCo)评估。结果显示,超越直觉方法在两个数据集上表现最优,静态3级光照下中位覆盖率0.789,SaCo 0.837;UBFC-rPPG上分别为0.826和0.917。同一参与者内,两种指标均未与心率误差、波形相关性或信噪比显著相关:252个相关系数中186个|ρ|<0.10,仅28个p<0.05(低于13个预期)。跨八种场景,仅超越直觉方法的覆盖率与性能指标相关(ρ=-0.43, +0.57, +0.43),其他注意力方法的SaCo反而反向变化。在40 lux光照下,该方法覆盖率降至0.180,SaCo降至-0.178,而运动干扰虽严重却未引发解释退化。
原文摘要 · Abstract (English)
Background. Remote photoplethysmography estimates the cardiovascular pulse from facial video, and its explanations have rested on inspecting heatmaps rather than on quantitative evidence about where a model reads it. We quantified the explanations and asked whether such explanations transfer between datasets and track model performance. Method. We trained eight condition-specific RhythmFormer models on NCKU-rPPG, recorded under three illumination levels, speaking, rotation, and cycling, estimated one heart rate per 5.12-second clip, and set them beside a UBFC-rPPG reproduction. Raw attention, rollout, attention flow, and Beyond Intuition were assessed by skin coverage and the Salience-guided Faithfulness Coefficient (SaCo). Results. Beyond Intuition ranked highest on both datasets, at median coverage 0.789 and SaCo 0.837 on Static level 3 against 0.826 and 0.917 on UBFC-rPPG; lower ranks differed. Within one participant of one condition, neither measure was related to a clip's heart-rate error, waveform correlation, or signal-to-noise ratio on either dataset: 186 of the 252 coefficients fell below $|\rho|=0.10$ and 28 reached $p<0.05$ against the 13 expected by chance. Across the eight scenarios only Beyond Intuition's coverage followed the three performance measures, at $\rho=-0.43$, $+0.57$, and $+0.43$, while the attention-only methods' SaCo ran opposite to each. It failed at 40 lux alone, its median coverage falling to 0.180 and its median SaCo to $-0.178$, whereas motion degraded the estimates far more without such a drop. Conclusions. Skin coverage and SaCo carry information complementary to the performance measures rather than a proxy for them: attributing to the skin does not guarantee an accurate estimate. What an attribution reveals about a condition is where the model looks rather than how faithfully its map is ordered.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。