arXiv:2512.11446cs.CV2025-12中稿 · the 33rd IEEE Inte…被引 1

构建帧级标注数据集,提升疲劳驾驶检测准确率

YawDD+: Frame-level Annotations for Accurate Yawn Prediction

  • 采用人机协同半自动化流程,实现视频帧级精确标注
  • 模型在边缘设备上达到99.34%分类准确率,检测mAP达95.69%
  • 适合需要实时、本地化疲劳检测的车载系统开发者

驾驶员疲劳是导致交通事故的主要原因,占事故总数的24%。打哈欠作为疲劳的早期行为信号,但现有方法因视频标注时间粒度粗糙导致系统性噪声而受限。为训练鲁棒的机器学习模型,需丰富监督标签以提取关键特征。本文提出一种人机协同的半自动化标注流程,将原YawDD数据集升级为帧级标注的YawDD+数据集,支持在NVIDIA Jetson NANO等边缘设备上高效训练与推理。基于此,在边缘平台训练的MNasNet分类器和YOLOv11检测器相比视频级标注,帧准确率提升最高达6%,mAP提升5%。在Jetson AGX上,MNasNet每轮训练仅需8.69分钟,推理速度高达115帧/秒,实现无需云端计算的实时疲劳监测。所提数据集与模型已公开。

原文摘要 · Abstract (English)

Driver fatigue remains a leading cause of road accidents, responsible for 24% of crashes. While yawning serves as an early behavioral indicator of fatigue, existing approaches face significant challenges due to the presence of systematic noise in video-annotated datasets arising from coarse temporal annotations. Training robust machine learning (ML) models requires rich supervisory labels that help learn salient features from the training data. Moreover, efficient on-device training and inference of models on edge devices is crucial in driver fatigue detection tasks to enable accurate real-time decisions on vehicles without reliance on cloud infrastructure. To address this issue, we develop a semi-automated labeling pipeline with human-in-the-loop verification to annotate YawDD videos to YawDD+ frame-level annotations, enabling more accurate model training on edge platforms such as NVIDIA Jetson NANO. Training the established MNasNet classifier and YOLOv11 detector architectures on YawDD+ improves frame accuracy by up to 6% and mAP by 5% over video-level supervision, achieving 99.34% classification accuracy and 95.69% detection mAP on Jetson NANO and AGX. Moreover, MNasNet completed the epoch time in just 8.69 min/epoch while delivering up to 115 frames-per-second (FPS) inference time on AGX, confirming that enhanced data quality alone supports on-device driver fatigue monitoring systems without server-side computation. The YawDD+ dataset and trained models are available online.

疲劳检测边缘计算数据标注视觉识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。