arXiv:2609.08117cs.CVcs.HC2026-09

构建首个连续司机动作预测大模型基准,提升驾驶意图预判能力。

DriveMotion: A Large-Scale Multi-Source Benchmark for Driver Motion Sequence Modeling and Forecasting

论文配图:DriveMotion: A Large-Scale Multi-Source Benchmark for Driver Motion Sequence Modeling and Forecasting
图 1 · 摘自论文原文
  • 融合多源数据构建393小时高精度司机动作序列,支持连续预测。
  • 动态锚定评估发现转向前手臂运动增加3.4倍,模型误差降低15%。
  • 适合自动驾驶、人机交互研究者,推动行为预测可复现性发展。

司机动作可反映当前行为、注意力及短期驾驶意图。然而,现有以司机为中心的数据集多聚焦于从短片段中识别预定义行为,而人体动作预测基准则主要关注车外运动。本文提出DriveMotion,一个用于连续司机动作预测的多源基准。该数据集包含360名驾驶员的393小时133关键点动作序列(10 Hz),整合自然驾驶数据、精选舱内视频与AIDE数据集,采用每关节有效性掩码和同步驾驶上下文统一表示。自然驾驶中存在长时间肢体活动受限,导致均匀采样评估受持续性主导,对短暂但行为意义显著的动作不敏感。为此,我们采用动态锚定评估:基于离线CAN信号识别车辆操作,并在操作前设置预测窗口,不向模型提供CAN信息。结果显示,操作前窗口内手臂运动是稳定驾驶对照组的3.4倍。在此锚定窗口上,学习模型相比持续性基线将预测误差降低最高达15%,且操作丰富训练使预测结果的部件状态F1值提升44%超过零运动参考。全多源数据集训练进一步使保留网页驾驶员上的预测误差比仅用BATON训练降低38%。DriveMotion提供身份无关划分、固定评估子集及参考实现,支持连续司机动作预测的可复现评估。数据集与基准已开放:https://huggingface.co/datasets/HenryYHW/DriveMotion。

原文摘要 · Abstract (English)

Driver motion can provide cues to ongoing behavior, attention, and near-term driving intent. However, most existing driver-centric datasets focus on recognizing predefined driver behaviors from short video clips, while human motion forecasting benchmarks largely target motion outside the vehicle. We introduce DriveMotion, a multi-source benchmark for continuous driver motion forecasting. DriveMotion contains 393 hours of 133-keypoint motion sequences at 10 Hz from 360 drivers, integrating naturalistic driving data, curated public in-cabin videos, and the AIDE dataset into a unified representation with per-joint validity masks and synchronized driving context. Naturalistic driving contains long periods of limited body movement, making uniformly sampled evaluation dominated by persistence and less sensitive to brief but behaviorally meaningful motion. To address this, we use dynamics-anchored evaluation, placing forecasting windows around vehicle maneuvers identified offline from CAN signals without providing CAN to the model at inference. Arm motion in pre-maneuver windows is 3.4x greater than in route-matched stable-driving controls. On these anchored windows, learned models reduce forecasting error over persistence by up to 15%, while maneuver-enriched training improves forecast-derived Part-State F1 by 44% over the zero-motion reference. Training on the full multi-source corpus further reduces forecasting error on held-out web drivers by 38% compared with BATON-only training. DriveMotion provides identity-disjoint splits, fixed evaluation subsets, and reference implementations for reproducible evaluation of continuous driver motion forecasting. The dataset and benchmark are available at https://huggingface.co/datasets/HenryYHW/DriveMotion

动作预测驾驶行为多源数据自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。