用人脸视频识酒驾,准确率超95%
Detection of Intoxicated Individuals from Facial Video Sequences via a Recurrent Fusion Model
- 融合图注意力网络与3D残差网络,动态加权特征
- 在3542段视频上达95.82%准确率,优于基线
- 适合公共安全场景的无感酒驾检测
饮酒是全球范围内的重大公共卫生问题,也是事故和死亡的主要原因。本文提出一种基于视频的人脸序列分析方法,用于酒精醉酒检测。该方法结合图注意力网络(GAT)进行面部关键点分析,以及使用3D ResNet提取时空视觉特征,并通过自适应优先级动态融合这些特征以提升分类性能。此外,我们构建了一个包含3,542个视频片段的专用数据集,源自202名个体。模型与两个基线方法(自定义3D-CNN和VGGFace+LSTM)进行了对比。实验结果表明,本方法在测试中达到95.82%的准确率、0.977的精确率和0.97的召回率,优于现有方法。研究结果证明了该模型在公共安全系统中实现非侵入式、可靠酒驾检测的潜力。
原文摘要 · Abstract (English)
Alcohol consumption is a significant public health concern and a major cause of accidents and fatalities worldwide. This study introduces a novel video-based facial sequence analysis approach dedicated to the detection of alcohol intoxication. The method integrates facial landmark analysis via a Graph Attention Network (GAT) with spatiotemporal visual features extracted using a 3D ResNet. These features are dynamically fused with adaptive prioritization to enhance classification performance. Additionally, we introduce a curated dataset comprising 3,542 video segments derived from 202 individuals to support training and evaluation. Our model is compared against two baselines: a custom 3D-CNN and a VGGFace+LSTM architecture. Experimental results show that our approach achieves 95.82% accuracy, 0.977 precision, and 0.97 recall, outperforming prior methods. The findings demonstrate the model's potential for practical deployment in public safety systems for non-invasive, reliable alcohol intoxication detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。