用统一离散符号表征视觉与动作,让自动驾驶规划更可控、可解释。
Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning

- 将视觉、动作、决策统一到离散令牌空间,实现多任务联合训练
- 在大规模驾驶数据上实现强规划性能,支持反事实分析和意外检测
- 分层决策+并行编辑,提升策略生成效率与可调控性
自动驾驶需推理自身行为如何影响未来世界演化,而非仅将观测映射为动作。现有端到端方法多依赖状态到动作的直接模仿,而世界模型常与下游策略生成弱对齐。本文提出Discrete-WAM,一种统一的离散视觉-动作世界-策略框架,将视觉观测、未来状态、高层决策与自车动作纳入共享令牌空间。基于此离散对齐,通过多任务、多阶段预训练联合训练世界建模、世界-策略建模与策略建模,使条件化未来预测直接支持策略生成。下游规划中,进一步将策略生成分解为层级决策预测与并行动作令牌编辑:决策令牌提供高层规划骨架,置信度调度高效优化密集未来动作。在大规模自动驾驶基准测试中,Discrete-WAM实现优异规划性能,同时支持可控未来生成、反事实评估、基于意外的世界模型分析及高效并行策略解码。结果表明,离散表示对齐、统一世界-策略训练与分层令牌编辑,为物理AI提供有前景的设计范式。
原文摘要 · Abstract (English)
Autonomous driving requires reasoning about how ego actions shape future world evolution, rather than merely mapping observations to actions. However, most end-to-end methods rely on direct state-to-action imitation, while existing world models often remain weakly aligned with downstream policy generation. We introduce Discrete-WAM, a unified discrete vision-action world-policy framework that represents visual observations, future states, high-level decisions, and ego actions within a shared token space. Built on this discrete alignment, Discrete-WAM jointly trains world modeling, world-policy modeling, and policy modeling through multi-task and multi-stage pretraining, allowing action-conditioned future prediction to directly support policy generation. For downstream planning, Discrete-WAM further decomposes policy generation into hierarchical decision prediction and parallel action-token editing, where the decision token provides a high-level planning skeleton and confidence-based scheduling refines dense future actions efficiently. Experiments on large-scale autonomous-driving benchmarks show that Discrete-WAM achieves strong planning performance while supporting controllable future generation, counterfactual evaluation, surprise-based world-model analysis, and efficient parallel policy decoding. These results suggest that discrete representation alignment, unified world-policy training, and hierarchical token editing provide a promising design paradigm for physical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。