自动分析行车视频,精准识别交叉口驾驶行为。
Data extraction and processing methods to aid the study of driving behaviors at intersections in naturalistic driving
- 用AI检测头部姿态和转向动作,量化驾驶员视线扫描。
- 94%准确识别转弯行为,定位误差仅1.1米以内。
- 适合交通行为研究者与智能驾驶系统开发者。
自然驾驶研究通过车载设备记录数月日常驾驶数据,因数据量大且类型多样,需自动化处理。本文介绍从车载记录系统中提取并分析交叉口驾驶员头部扫描的方法,该系统记录了车速、GPS位置、场景视频和舱内视频。开发了定制工具用于标记交叉口,同步位置与视频数据,并裁剪交叉口前后100米的舱内与场景视频。采用自研宽视角头部姿态检测AI模型分析舱内视频,计算水平方向大于20度的头部扫描。场景视频使用YOLO目标检测模型识别交通灯、停车标志、行人及其它车辆。通过车辆自运动模式独立检测转弯动作。利用车辆移动时路面亮度变化模式检测停止线。结合场景视频信息与车速数据,设计规则算法推断交叉口类型、行驶动作及边界范围。共处理来自马萨诸塞州与加利福尼亚州3辆汽车在城市和郊区行驶的190个交叉口数据。自动化视频处理算法对交叉口标志与转弯动作的检测准确率分别为100%和94%。车辆进入交叉口的定位中位误差为1.1[0.4–4.9]米,时间误差为0.2[0.1–0.54]秒;与真实边界重叠中位数为0.88[0.82–0.93]。
原文摘要 · Abstract (English)
Naturalistic driving studies use devices in participants' own vehicles to record daily driving over many months. Due to diverse and extensive amounts of data recorded, automated processing is necessary. This report describes methods to extract and characterize driver head scans at intersections from data collected from an in-car recording system that logged vehicle speed, GPS location, scene videos, and cabin videos. Custom tools were developed to mark the intersections, synchronize location and video data, and clip the cabin and scene videos for +/-100 meters from the intersection location. A custom-developed head pose detection AI model for wide angle head turns was run on the cabin videos to estimate the driver head pose, from which head scans >20 deg were computed in the horizontal direction. The scene videos were processed using a YOLO object detection model to detect traffic lights, stop signs, pedestrians, and other vehicles on the road. Turning maneuvers were independently detected using vehicle self-motion patterns. Stop lines on the road surface were detected using changing intensity patterns over time as the vehicle moved. The information obtained from processing the scene videos, along with the speed data was used in a rule-based algorithm to infer the intersection type, maneuver, and bounds. We processed 190 intersections from 3 vehicles driven in cities and suburban areas from Massachusetts and California. The automated video processing algorithm correctly detected intersection signage and maneuvers in 100% and 94% of instances, respectively. The median [IQR] error in detecting vehicle entry into the intersection was 1.1[0.4-4.9] meters and 0.2[0.1-0.54] seconds. The median overlap between ground truth and estimated intersection bounds was 0.88[0.82-0.93].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。