用生理信号指导视频情绪识别,提升准确率且无需佩戴传感器。
BioKD: Selective Physiology-to-Video Knowledge Distillation via Reliability Gate for Emotion Recognition

- 通过可靠性门控机制筛选可靠生理信号,动态调节知识迁移强度。
- 在DEAP数据集上达68.01%(试次级)和65.29%(个体级)准确率。
- 无需额外计算开销,适合真实场景部署的轻量级情绪识别系统。
为解决视频情绪识别在行为线索模糊或社会掩饰下的局限性,以及生理信号难以部署的问题,本文提出一种可靠性感知的生理信号到视频的知识蒸馏框架BioKD。该框架在训练阶段利用生理信号作为特权信息,引导视频学生模型学习深层情感表征,推理时仅依赖非侵入式视频输入。针对生理信号因个体差异、信号伪影和时间不一致带来的高噪声与不稳定性,BioKD引入样本级可靠性感知门控机制与渐进式蒸馏策略,自适应调节知识传递强度,抑制不可靠监督导致的负向迁移,实现更稳定的跨模态蒸馏。在DEAP和AMIGOS数据集上的实验表明,BioKD在试次级与个体级评估协议下均优于主流基线。例如,在更具有挑战性的个体级设置下,其在DEAP数据集上对唤醒度的识别准确率达65.29%,显著优于熵权重策略,验证了显式建模监督可靠性的重要性。此外,BioKD不增加推理时开销,且无需生理传感与多模态同步。
原文摘要 · Abstract (English)
To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD. The proposed framework leverages physiological signals as privileged information during training to guide a video-based student model in learning deep affective representations, while relying solely on non-intrusive video inputs at inference time. To cope with the high noise and instability of physiological teacher supervision caused by inter-subject variability, signal artifacts, and temporal inconsistency, BioKD incorporates a sample-wise reliability-aware gating mechanism together with a progressive distillation strategy. By adaptively regulating the strength of knowledge transfer, the framework suppresses negative transfer induced by unreliable physiological supervision and enables more stable cross-modal distillation. Experiments on DEAP and AMIGOS show that BioKD consistently outperforms representative baselines under both trial-wise and subject-wise evaluation protocols for valence and arousal recognition. For example, BioKD achieves 68.01\% on DEAP (trial-wise arousal) and 65.29\% under the more challenging subject-wise setting, demonstrating improved performance under a subject-independent evaluation setting. Further analyses show that BioKD effectively mitigates overconfident teacher errors and outperforms an entropy-only weighting strategy, confirming the importance of explicitly modeling supervision reliability. In addition, BioKD introduces no additional inference-time overhead relative to the same video student architecture and removes the need for physiological sensing and multimodal synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。