用语言提示微调自动驾驶轨迹,尤其在指令不可靠时效果显著
NudgeVAD: Language-Nudged End-to-End Driving via FiLM Residuals

- 通过语言条件残差模块,在冻结规划器基础上实现可控微调
- 指令不可靠时,语言使平均位移误差降低至2.806米,优于无语言模型
- 适合需要语言控制但命令不稳定的自动驾驶场景
自然语言指令有望提升端到端自动驾驶的可控性,但当规划器已有可靠高层指令时,其优势可能被掩盖。我们提出NudgeVAD,一种基于冻结规划器的残差框架,利用语言作为对VAD轨迹的校准提示。通过身份初始化的FiLM和零初始化的残差头,NudgeVAD在初始化时等同于冻结规划器,学习到的偏差仅来自语言条件残差。我们在指令可靠性轴上评估NudgeVAD:在可靠指令下,语言虽能改善初始规划器,但与计算量相当的无语言微调模型VAD-FT(UNCOND)相比几乎冗余;而在随机指令下,移除文本导致ADE6升至3.166米,而NudgeVAD保留文本可恢复至2.806米,且优于VAD-FT(UNCOND)0.312米。结果表明语言并非普遍增益,仅在类别指令通道不可靠时最为关键。
原文摘要 · Abstract (English)
Natural-language instructions promise controllable end-to-end driving, but their benefit can be hidden when planners already receive reliable high-level commands. We propose NudgeVAD, a frozen-planner residual framework that uses language as a calibrated nudge to a VAD trajectory. With identity-initialized FiLM and a zero-initialized residual head, NudgeVAD is equivalent to the frozen planner at initialization, so learned deviations arise only from language-conditioned residuals. We evaluate NudgeVAD along a command-reliability axis. With reliable commands, language improves the initial planner but becomes nearly redundant once compared against VAD-FT (UNCOND), a compute-matched VAD model fine-tuned without language. With random commands, however, language becomes essential: detaching text degrades ADE6s to 3.166 m, while NudgeVAD with text recovers 2.806 m and outperforms VAD-FT (UNCOND) by 0.312 m. These results show that language is not universally additive; it is most valuable when the categorical command channel is unreliable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。