让世界模型直接指导规划,提升自动驾驶决策可靠性
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction

- 构建统一架构,通过无动作预测未来状态来辅助规划
- 仅用前视摄像头实现媲美多视角方法的预测性能
- 适合自动驾驶、智能机器人等需前瞻决策的场景
尽管驾驶世界模型取得显著进展,其在自主系统中的潜力仍待挖掘:现有模型多用于世界模拟,与轨迹规划分离。近期工作尝试统一建模与规划,但二者协同机制尚不清晰。本文提出政策世界模型(PWM),将世界建模与轨迹规划整合于统一框架,并通过无需动作的未来状态预测机制,使世界知识反哺规划。通过协同状态-动作预测,PWM可模拟人类前瞻感知,提升规划可靠性。为提升视频预测效率,引入动态增强的并行标记生成机制,包含上下文引导分词器和自适应动态焦点损失。仅使用前视摄像头输入,性能即达或超越依赖多视角、多模态输入的先进方法。代码与模型权重将在 https://github.com/6550Zhao/Policy-World-Model 发布。
原文摘要 · Abstract (English)
Despite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and decoupled from trajectory planning. While recent efforts aim to unify world modeling and planning in a single framework, the synergistic facilitation mechanism of world modeling for planning still requires further exploration. In this work, we introduce a new driving paradigm named Policy World Model (PWM), which not only integrates world modeling and trajectory planning within a unified architecture, but is also able to benefit planning using the learned world knowledge through the proposed action-free future state forecasting scheme. Through collaborative state-action prediction, PWM can mimic the human-like anticipatory perception, yielding more reliable planning performance. To facilitate the efficiency of video forecasting, we further introduce a dynamically enhanced parallel token generation mechanism, equipped with a context-guided tokenizer and an adaptive dynamic focal loss. Despite utilizing only front camera input, our method matches or exceeds state-of-the-art approaches that rely on multi-view and multi-modal inputs. Code and model weights will be released at https://github.com/6550Zhao/Policy-World-Model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。