用动态符号压缩未来交通变化,让自动驾驶决策更精准高效。
DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving
- 先预测世界动态符号,再生成驾驶动作,提升决策合理性
- 在三个数据集上优于文本和视觉思维链方法,成功率显著提升
- 适合需要高可靠性的自动驾驶决策系统研究者参考
我们提出 DynVLA,一种面向自动驾驶的视觉-语言-动作模型,引入新型思维链范式——动态思维链(Dynamics CoT)。DynVLA 在生成动作前预测紧凑的世界动态演化,实现更具物理合理性的决策。为获得紧凑的动态表征,模型引入动态标记器(Dynamics Tokenizer),将未来演化压缩为少量动态标记。针对交互密集场景中的丰富环境动态,模型解耦自车中心与环境中心动态,提升建模精度。通过监督微调(SFT)与强化反馈训练(RFT),DynVLA 学会生成动态标记以指导动作,同时保持低延迟推理。相比缺乏细粒度时空理解的文本思维链,以及因密集图像预测带来冗余的视觉思维链,动态思维链以紧凑、可解释且高效的形式捕捉世界演化。在 NAVSIM、Bench2Drive 及大规模自研数据集上的实验表明,DynVLA 持续优于文本与视觉思维链方法,验证了动态思维链的有效性与实用价值。
原文摘要 · Abstract (English)
We propose DynVLA, a driving VLA model that introduces a new CoT paradigm termed Dynamics CoT. DynVLA forecasts compact world dynamics before action generation, enabling more informed and physically grounded decision-making. To obtain compact dynamics representations, DynVLA introduces a Dynamics Tokenizer that compresses future evolution into a small set of dynamics tokens. Considering the rich environment dynamics in interaction-intensive driving scenarios, DynVLA decouples ego-centric and environment-centric dynamics, yielding more accurate world dynamics modeling. We then train DynVLA to generate dynamics tokens before actions through SFT and RFT, improving decision quality while maintaining latency-efficient inference. Compared to Textual CoT, which lacks fine-grained spatiotemporal understanding, and Visual CoT, which introduces substantial redundancy due to dense image prediction, Dynamics CoT captures the evolution of the world in a compact, interpretable, and efficient form. Extensive experiments on NAVSIM, Bench2Drive, and a large-scale in-house dataset demonstrate that DynVLA consistently outperforms Textual CoT and Visual CoT methods, validating the effectiveness and practical value of Dynamics CoT. Project Page: https://yaoyao-jpg.github.io/dynvla.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。