arXiv:2609.03225cs.RO2026-09

提出可长期保持一致、感知交互的驾驶世界模型,支持多种驾驶风格。

Long-Horizon Consistent and Interaction-Aware World Models for Multi-Style End-to-End Driving

论文配图:Long-Horizon Consistent and Interaction-Aware World Models for Multi-Style End-to-End Driving
图 1 · 摘自论文原文
  • 用门控交叉注意力融合历史隐状态,稳定长时想象推演
  • 分离自车相关与无关交互状态,提升复杂路况决策效率
  • 基于轨迹相对优势优化,无需重训即可适配不同驾驶风格

端到端自动驾驶正越来越多地采用基于世界模型的强化学习框架,通过想象式回放提高学习效率。然而现有世界模型存在三大缺陷:长时程想象推演的时间不一致性、对自车-环境交互建模不足、难以适应多样驾驶风格。为此,我们提出StyleDrive框架,联合实现长时一致性、显式解耦交互状态、统一范式下多风格策略优化。首先引入时间一致性正则化,通过门控交叉注意力融合历史隐状态,稳定长时想象推演并缓解误差累积。其次设计显式状态解耦模块,将自车相关与无关交互状态分离,提升复杂交通场景下决策的可解释性与效率。第三,通过组相对策略优化(Group Relative Policy Optimization),以轨迹级相对优势替代逐步奖励优化,降低奖励方差,支持多样化驾驶风格而无需重训练。我们在Bench2Drive闭环驾驶基准上评估,驾驶得分88.44(较前最佳方法提升17.08),成功率66.82(提升16.58)。此外,将StyleDrive部署于真实自动导引车平台,在动态驾驶场景中展示出良好的仿真到现实迁移能力。

原文摘要 · Abstract (English)

End-to-end autonomous driving has increasingly adopted world model-based reinforcement learning frameworks to improve learning efficiency through \textit{imagined rollouts}. However, existing world models suffer from three key limitations: temporal inconsistency in long-horizon imagined rollouts, inadequate modeling of ego-environment interactions, and limited adaptability to diverse driving styles. To address these challenges, we propose \textit{StyleDrive}, a world-model-based learning framework that jointly enforces long-horizon consistency, explicitly disentangles interactive traffic states, and supports multi-style policy optimization within a unified learning paradigm. First, we introduce a temporal consistency regularization that integrates historical latent states through gated cross-attention, stabilizing long-horizon imagined rollouts and mitigating error accumulation. Second, we design an explicit state disentanglement module that separates ego-relevant from ego-irrelevant interactive states, enabling more interpretable and efficient decision-making in complex traffic scenarios. Third, we enable multi-style driving behaviors through Group Relative Policy Optimization, which replaces per-step reward optimization with trajectory-wise relative advantages, reducing reward variance and supporting diverse driving styles without retraining. We evaluate StyleDrive on the Bench2Drive closed-loop driving benchmark, achieving a driving score of 88.44 (+17.08 over the previous best world model-based method) and a success rate of 66.82 (+16.58). Furthermore, we deploy StyleDrive on a real automated guided vehicle platform and demonstrate promising sim-to-real transfer capability in dynamic driving scenarios.

自动驾驶世界模型多风格驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。