对比单视角与双视角视频,发现融合环境信息能提升注意力检测,但需适配模型架构。
A Contextual Analysis of Driver-Facing and Dual-View Video Inputs for Distraction Detection in Naturalistic Driving Environments
- 用双摄像头同步采集真实驾驶数据,测试三种动作识别模型在单/双视图下的表现
- 慢速仅路径模型(SlowOnly)双视图下准确率提升9.8%,而双路径模型(SlowFast)下降7.2%
- 强调多视角融合需架构支持,否则易产生表征冲突,适合做驾驶监控系统设计参考
尽管基于计算机视觉的分心驾驶检测日益受到关注,现有模型大多仅依赖驾驶员视角,忽视了影响驾驶行为的关键环境上下文。本研究探究在自然驾驶环境中,结合道路前方视角与驾驶员视角是否能提升分心检测精度。利用真实驾驶场景中同步采集的双摄像头视频,我们对三种主流时空动作识别架构(SlowFast-R50、X3D-M、SlowOnly-R50)进行评估,分别在仅驾驶员视角和堆叠双视角输入两种配置下测试。结果显示,虽然上下文信息在部分模型中可提升性能,但增益高度依赖于模型结构。单路径的SlowOnly模型在双视角输入下准确率提升9.8%,而双路径的SlowFast模型反而下降7.2%,归因于表征冲突。结果表明,单纯增加视觉上下文并不足够,若模型未专门设计支持多视角融合,反而会引入干扰。本研究是首个使用真实驾驶数据系统比较单/双视图分心检测模型的工作,强调未来多模态驾驶员监控系统需采用融合感知的设计范式。
原文摘要 · Abstract (English)
Despite increasing interest in computer vision-based distracted driving detection, most existing models rely exclusively on driver-facing views and overlook crucial environmental context that influences driving behavior. This study investigates whether incorporating road-facing views alongside driver-facing footage improves distraction detection accuracy in naturalistic driving conditions. Using synchronized dual-camera recordings from real-world driving, we benchmark three leading spatiotemporal action recognition architectures: SlowFast-R50, X3D-M, and SlowOnly-R50. Each model is evaluated under two input configurations: driver-only and stacked dual-view. Results show that while contextual inputs can improve detection in certain models, performance gains depend strongly on the underlying architecture. The single-pathway SlowOnly model achieved a 9.8 percent improvement with dual-view inputs, while the dual-pathway SlowFast model experienced a 7.2 percent drop in accuracy due to representational conflicts. These findings suggest that simply adding visual context is not sufficient and may lead to interference unless the architecture is specifically designed to support multi-view integration. This study presents one of the first systematic comparisons of single- and dual-view distraction detection models using naturalistic driving data and underscores the importance of fusion-aware design for future multimodal driver monitoring systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。