建模司机对路况的反应,实现车内动态长期预测
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

- 用双流结构分别编码路况和司机状态,通过门控机制因果耦合
- 在AIDE数据集上实现长时序高动态反应预测,效果优于基线
- 适合自动驾驶共享控制场景中的安全干预与人机协同研究
安全的L2/L3级自动驾驶需预判人机共控切换时驾驶员的反应。现有世界模型多聚焦外部环境预测,而车内智能仍以识别为主,缺乏对司机行为的多步滚动预测能力。本文提出Driver-WM,一种以司机为中心的潜在世界模型,基于外部交通情境因果地滚动预测车内动态。该模型将物理运动预测与辅助的行为、情绪语义识别统一建模。在冻结的视觉-语言特征构建的紧凑潜在空间中,采用双流架构分别编码外部交通与内部司机状态,并通过门控因果注入机制定向耦合,利用学习到的向量门控调节外部上下文扰动,严格保证时间因果性。在AIDE数据集上的实验表明,该模型在高动态反应片段上实现了稳健的长时序预测,提升了司机与交通语义对齐度,并能可控干预以揭示外部到内部的作用机制。
原文摘要 · Abstract (English)
Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models forecast the external environment, in-cabin intelligence remains strictly recognition-oriented and lacks multi-step rollout capabilities for driver dynamics. We introduce Driver-WM, a driver-centric latent world model that rolls out in-cabin dynamics causally conditioned on out-cabin traffic context. This formulation unifies physical kinematics forecasting with auxiliary behavioral and emotional semantic recognition. Operating in a compact latent space constructed from frozen vision-language features, Driver-WM adopts a dual-stream architecture to separately encode external traffic and internal driver states. These streams are directionally coupled via a gated causal injection mechanism, which uses a learned vector gate to modulate external contextual perturbations while strictly enforcing temporal causality. Experiments on AIDE show robust long-horizon forecasting on reactive high-motion clips, improved driver/traffic semantic alignment, and controlled interventions that expose the external-to-internal mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。