用双循环自编码器联合建模人脸表情时空特征,提升驾驶状态检测准确率。
Spatiotemporal Facial Action Unit Detection using Twin Cycle Autoencoders for Driver Monitoring
- 设计双分支循环自编码结构,分别处理外观与动态,实现跨时序一致性约束。
- 在DISFA、BP4D及真实驾驶数据集上,对微弱和快速变化的面部动作单元检测提升显著。
- 模型可在嵌入式设备实时运行,适合部署于量产级智能驾驶辅助系统。
驾驶监控系统日益依赖面部线索实时推断疲劳、分心与认知负荷。面部动作单元(AUs)基于面部动作编码系统(FACS),提供客观可解释的状态表征,但其在驾驶场景中的自动检测受光照不足、遮挡、头部姿态变化以及动作强度低、持续时间短等因素影响。现有检测方法多将空间外观与时间动态分开处理,难以利用大量未标注驾驶视频中的自监督信号。本文提出孪生循环自编码器(TCA),由两个耦合的循环一致自编码分支组成:空间循环自编码器通过图像级循环一致性分离动作相关外观与身份信息;时间循环自编码器通过对潜空间动作轨迹施加前向-后向一致性,捕捉起始-峰值-结束动态。两分支通过跨分支潜空间对齐损失耦合,并经注意力模块融合后进行多标签动作单元分类。在DISFA、BP4D及舱内自然驾驶数据集上评估显示,相比CNN-RNN、3D-CNN与图结构基线,TCA在低强度与快速切换的动作单元(如疲劳相关AU45、AU43,打哈欠相关AU26)上均有稳定提升。进一步验证表明,模型可在嵌入式Jetson Xavier NX平台实现实时推理,支持生产级高级驾驶辅助系统(ADAS)应用。
原文摘要 · Abstract (English)
Driver monitoring systems (DMS) increasingly rely on facial cues to infer drowsiness, distraction, and cognitive load in real time. Facial Action Units (AUs), grounded in the Facial Action Coding System (FACS), provide an objective and interpretable representation of such states, but their automatic detection in the driving context is complicated by low and variable illumination, partial occlusion, head-pose variation, and the subtlety and short duration of relevant AU activations. Existing AU detectors largely treat spatial appearance and temporal dynamics separately, limiting their ability to exploit self-supervisory signal from abundant unlabeled driving video. We propose the Twin Cycle Autoencoder (TCA), a spatiotemporal architecture composed of two coupled cycle-consistent autoencoder branches: a Spatial Cycle Autoencoder that disentangles AU-relevant appearance from identity through image-level cycle consistency, and a Temporal Cycle Autoencoder that enforces forward-backward consistency over latent AU trajectories to capture onset-apex-offset dynamics. The two branches are coupled through a cross-branch latent alignment loss and fused via an attention module before multi-label AU classification. We evaluate TCA on the DISFA and BP4D benchmarks and on an in-cabin naturalistic driving dataset, and observe consistent improvements over CNN-RNN, 3D-CNN, and graph-based AU baselines, particularly for low-intensity and rapidly transitioning AUs relevant to fatigue (AU45, AU43) and yawning (AU26). We further show the model sustains real-time throughput on an embedded Jetson Xavier NX platform, supporting its use in production-grade advanced driver assistance systems (ADAS).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。