构建首个双向驾驶自动化切换多模态数据集,助力智能座舱交互设计
BATON: A Multimodal Benchmark for Bidirectional Automation Transition Observation in Naturalistic Driving
- 采集127名司机136.6小时真实道路数据,同步车外视频、车内视频、CAN总线等多模态信号
- 发现仅用外部视频无法可靠预测控制交接,融合车辆状态与路线信息性能显著提升
- 揭示接管比交出更渐进,适合长时预警;交出依赖即时线索,需快速响应
现有量产车型的驾驶自动化系统要求驾驶员自主决定何时启用并持续保持警觉,这带来巨大认知负荷,导致学习成本高、体验差且存在过度依赖或接管延迟的安全风险。准确预测驾驶员何时交出控制权和何时收回控制权,对设计主动式、情境感知的人机交互至关重要。然而现有数据集很少同时包含道路场景、驾驶员状态、车辆动态和路线环境等多模态上下文。为此,我们提出BATON,一个大规模自然驾驶数据集,涵盖127名驾驶员和136.6小时实际驾驶数据。该数据集同步记录前视视频、车内视频、解码后的CAN总线信号、基于雷达的前车交互信息及GPS获取的路线上下文,形成每个控制转换环节的闭环多模态记录。我们定义三项基准任务:驾驶行为理解、交出控制预测和接管控制预测,并评估了从序列模型到经典分类器再到零样本视觉语言模型的基线方法。结果表明,仅使用视觉输入不足以实现可靠的过渡预测:前视视频捕捉道路环境但忽略驾驶员状态,而车内视频反映驾驶员准备度却无法体现外部场景。结合CAN信号与路线上下文可显著提升性能,证明多模态之间具有强互补性。进一步发现,接管事件发展更缓慢,适合较长预测窗口;而交出事件更依赖即时上下文线索,揭示出直接指导辅助驾驶人机设计的不对称性。
原文摘要 · Abstract (English)
Existing driving automation (DA) systems on production vehicles rely on human drivers to decide when to engage DA while requiring them to remain continuously attentive and ready to intervene. This design demands substantial situational judgment and imposes significant cognitive load, leading to steep learning curves, suboptimal user experience, and safety risks from both over-reliance and delayed takeover. Predicting when drivers hand over control to DA and when they take it back is therefore critical for designing proactive, context-aware HMI, yet existing datasets rarely capture the multimodal context, including road scene, driver state, vehicle dynamics, and route environment. To fill this gap, we introduce BATON, a large-scale naturalistic dataset capturing real-world DA usage across 127 drivers, and 136.6 hours of driving. The dataset synchronizes front-view video, in-cabin video, decoded CAN bus signals, radar-based lead-vehicle interaction, and GPS-derived route context, forming a closed-loop multimodal record around each control transition. We define three benchmark tasks: driving action understanding, handover prediction, and takeover prediction, and evaluate baselines spanning sequence models, classical classifiers, and zero-shot VLMs. Results show that visual input alone is insufficient for reliable transition prediction: front-view video captures road context but not driver state, while in-cabin video reflects driver readiness but not the external scene. Incorporating CAN and route-context signals substantially improves performance over video-only settings, indicating strong complementarity across modalities. We further find takeover events develop more gradually and benefit from longer prediction horizons, whereas handover events depend more on immediate contextual cues, revealing an asymmetry with direct implications for HMI design in assisted driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。