arXiv:2603.09086cs.ROcs.AI2026-03被引 2

提出统一的潜在空间框架,整合自动驾驶世界模型最新进展。

Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges

  • 按潜在表示的目标与形式分类,构建系统化设计空间。
  • 提出五项核心机制,提升模型鲁棒性与可部署性。
  • 提供闭环评估指标,助力模型验证与资源优化。

生成式世界模型与视觉-语言-动作(VLA)系统正快速重塑自动驾驶,实现可扩展仿真、长时序预测及能力丰富的决策。其中,潜在表示作为核心计算基础:压缩多传感器高维观测,支持时间一致的滚动推演,并为规划、推理与可控生成提供接口。本文提出统一的潜在空间框架,整合自动驾驶领域世界模型的最新进展。该框架基于潜在表示的目标与形式(潜在世界、潜在动作、潜在生成器;连续状态、离散标记、混合形式)以及几何、拓扑与语义结构先验进行组织。在此分类基础上,提出五个跨领域核心机制(结构同构性、长时序时间稳定性、语义与推理对齐、价值对齐目标与后训练、自适应计算与深思),并将其与模型鲁棒性、泛化性和可部署性关联。此外,本文设计了具体的评估方案,包括闭环度量套件和资源感知的深思成本,以缓解开环与闭环之间的差距。最后,识别出推动决策就绪、可验证、资源高效自动驾驶的可行动研究方向。

原文摘要 · Abstract (English)

Emerging generative world models and vision-language-action (VLA) systems are rapidly reshaping automated driving by enabling scalable simulation, long-horizon forecasting, and capability-rich decision making. Across these directions, latent representations serve as the central computational substrate: they compress high-dimensional multi-sensor observations, enable temporally coherent rollouts, and provide interfaces for planning, reasoning, and controllable generation. This paper proposes a unifying latent-space framework that synthesizes recent progress in world models for automated driving. The framework organizes the design space by the target and form of latent representations (latent worlds, latent actions, latent generators; continuous states, discrete tokens, and hybrids) and by structural priors for geometry, topology, and semantics. Building on this taxonomy, the paper articulates five cross-cutting internal mechanics (i.e, structural isomorphism, long-horizon temporal stability, semantic and reasoning alignment, value-aligned objectives and post-training, as well as adaptive computation and deliberation) and connects these design choices to robustness, generalization, and deployability. The work also proposes concrete evaluation prescriptions, including a closed-loop metric suite and a resource-aware deliberation cost, designed to reduce the open-loop / closed-loop mismatch. Finally, the paper identifies actionable research directions toward advancing latent world model for decision-ready, verifiable, and resource-efficient automated driving.

自动驾驶世界模型潜在空间评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。