arXiv:2603.07264cs.ROcs.AI2026-03

将车辆运动规律融入模型,提升自动驾驶数据效率

Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving

  • 用运动学信息增强观测编码,让潜空间动态更符合物理规律
  • 在仿真环境中显著提升样本效率和驾驶表现,优于基线方法
  • 适合追求高效安全训练的自动驾驶算法研究者

由于大规模真实交互成本高且存在安全风险,数据高效学习仍是自动驾驶的核心挑战。尽管基于世界模型的强化学习可通过潜空间想象优化策略,但现有方法常缺乏对驾驶任务至关重要的空间与运动学结构的显式建模。本文在递归状态空间模型(RSSM)基础上,提出一种面向自动驾驶的运动学感知潜世界模型框架。通过在观测编码器中引入车辆运动学信息,使潜空间转移过程符合物理运动规律;同时,利用几何感知监督使RSSM潜状态捕捉任务相关的空间结构,而不仅限于像素重建。所提结构化潜动态提升了长时程想象的保真度,并稳定了策略优化。在驾驶仿真基准测试中,该方法在样本效率与驾驶性能上均持续优于模型无关及基于像素的世界模型基线。消融实验进一步验证了该设计提升了潜空间中的空间表征质量。结果表明,将运动学约束融入基于RSSM的世界模型,为自动驾驶策略学习提供了一种可扩展且物理合理的范式。

原文摘要 · Abstract (English)

Data-efficient learning remains a central challenge in autonomous driving due to the high cost and safety risks of large-scale real-world interaction. Although world-model-based reinforcement learning enables policy optimization through latent imagination, existing approaches often lack explicit mechanisms to encode spatial and kinematic structure essential for driving tasks. In this work, we build upon the Recurrent State-Space Model (RSSM) and propose a kinematics-aware latent world model framework for autonomous driving. Vehicle kinematic information is incorporated into the observation encoder to ground latent transitions in physically meaningful motion dynamics, while geometry-aware supervision regularizes the RSSM latent state to capture task-relevant spatial structure beyond pixel reconstruction. The resulting structured latent dynamics improve long-horizon imagination fidelity and stabilize policy optimization. Experiments in a driving simulation benchmark demonstrate consistent gains over both model-free and pixel-based world-model baselines in terms of sample efficiency and driving performance. Ablation studies further verify that the proposed design enhances spatial representation quality within the latent space. These results suggest that integrating kinematic grounding into RSSM-based world models provides a scalable and physically grounded paradigm for autonomous driving policy learning.

自动驾驶世界模型运动学数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。