用RGB视频学会识别疼痛,即使缺失热成像和深度信息也能准判。
ReMiX-MAE: Learning Missing-Channel Cross-Modal Representations from RGB-Only Clinical Facial Videos for Sympathetic-Mediated Pain Assessment

- 通过自监督学习从多模态视频中重建缺失的热/深度信号,增强鲁棒性。
- 在五分类疼痛评估中,伪多模态特征比纯RGB提升性能,且跨数据集泛化更强。
- 适合医疗场景下标注少、设备受限的疼痛自动评估任务。
真实临床环境中的自动疼痛评估受限于缺乏带弱标签(通常为序列级自评)的临床面颜视频数据,且疼痛线索在RGB图像中常表现细微或接近中性,而热成像与深度信号虽具信息量但难以常规部署。为此,我们提出ReMiX-MAE(Reconstructing Missing Channel Cross-Modal Masked Autoencoder),一种自监督多模态掩码预训练框架,能从同步的RGB、热成像和深度视频中学习可迁移的面部表征,并显式训练对模态缺失的鲁棒性,支持仅使用RGB的部署。为填补临床面颜疼痛数据中视频级自评与纵向治疗轨迹的空白,我们构建了包含多次就诊前后配对记录的“自主神经介导疼痛”(SMP)数据集。在仅使用RGB的部署条件下,我们通过直接特征提取与从RGB解码出的伪多模态特征评估ReMiX-MAE。结果表明,其在SMP数据集上持续优于仅基于RGB的掩码自编码器基线,尤其在困难的五分类任务中,伪多模态特征带来额外增益。在外部数据集上,ReMiX-MAE展现出更强的鲁棒性与标签效率,凸显其在数据稀缺的临床场景中的优势。
原文摘要 · Abstract (English)
Automated pain assessment in real clinics is limited by scarce clinically grounded facial video data with weak labels (often sequence-level self-report) and by the fact that pain cues can be subtle or near-neutral in RGB, while thermal and depth signals are informative yet impractical to deploy routinely. To address these challenges, we propose ReMiX-MAE (Reconstructing Missing Channel Cross-Modal Masked Autoencoder), a self-supervised multimodal masked pretraining framework that learns transferable facial representations from synchronized RGB, thermal, and depth videos and explicitly trains robustness to missing modalities, enabling RGB-only deployment. To fill the gap of clinically grounded facial pain data with video-level self-report and longitudinal treatment trajectories, we collect the Sympathetic Mediated Pain (SMP) dataset with paired pre- and post-recordings across multiple visits. Under RGB-only deployment, we evaluate ReMiX-MAE using both direct feature extraction and pseudo-multimodal features decoded from RGB. ReMiX-MAE consistently outperforms an RGB-only masked autoencoder baseline on SMP, with pseudo-multimodal features providing additional gains in the challenging five-class setting. Across external datasets, ReMiX-MAE further shows more robust and label-efficient transfer than RGB-only baselines, highlighting its advantage in data-limited clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。