融合人脸视频与生理信号,提升驾驶时压力识别准确率。
Combining Facial Videos and Biosignals for Stress Estimation During Driving
- 用3D人脸模型提取56维动态特征,捕捉细微表情与头部姿态变化。
- 跨模态注意力融合使准确率从51.0%升至86.7%,AUROC达92.0%。
- 适用于生理信号采集困难场景,适合智能驾驶与健康监测应用。
可靠的压力识别在医疗监测和驾驶等安全关键系统中至关重要。尽管常通过皮电反应和心率等生理信号检测压力,面部活动可无感获取,提供互补信息。本文提出一种多模态压力估计框架,结合面部视频与生理信号,在生理信号采集困难时仍有效。采用密集3D形态模型表示面部行为,生成56维描述符,捕捉随时间变化的细微表情与头部姿态动态。通过对比基线与压力诱导阶段的配对假设检验发现,56个面部成分中有38个表现出与生理指标相当的一致性、阶段特异性压力响应。基于此,构建基于Transformer的时序建模框架,评估单模态、早期融合与跨模态注意力策略。跨模态注意力融合3D面部特征与生理信号显著优于仅使用生理信号,将AUROC从52.7%提升至92.0%,准确率从51.0%提升至86.7%。虽在驾驶数据上验证,该框架与流程可推广至其他压力识别场景。
原文摘要 · Abstract (English)
Reliable stress recognition is critical in applications such as medical monitoring and safety-critical systems, including real-world driving. While stress is commonly detected using physiological signals such as perinasal perspiration and heart rate, facial activity provides complementary cues that can be captured unobtrusively from video. We propose a multimodal stress estimation framework that combines facial videos and physiological signals, remaining effective even when biosignal acquisition is challenging. Facial behavior is represented using a dense 3D Morphable Model, yielding a 56-dimensional descriptor that captures subtle expression and head-pose dynamics over time. To study how stress modulates facial motion, we perform extensive experiments alongside established physiological markers. Paired hypothesis tests between baseline and stressor phases show that 38 of 56 facial components exhibit consistent, phase-specific stress responses comparable to physiological markers. Building on these findings, we introduce a Transformer-based temporal modeling framework and evaluate unimodal, early-fusion, and cross-modal attention strategies. Cross-modal attention fusion of 3D-derived facial features with physiological signals substantially improves performance over physiological signals alone, increasing AUROC from 52.7% and accuracy from 51.0% to 92.0% and 86.7%, respectively. Although evaluated on driving data, the proposed framework and protocol may generalize to other stress estimation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。