arXiv:2601.12672cs.CV2026-01AAAI被引 6

用视觉语言模型直接生成驾驶挑战场景,提升自动驾驶系统鲁棒性。

VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness

  • 将视觉语言模型融入闭环训练,直接编辑周边车辆轨迹
  • 在长尾罕见场景下使自动驾驶策略安全率显著提升
  • 适合研究自动驾驶安全增强与智能对抗生成的学者

自动驾驶系统的安全部署受制于长尾问题——真实数据中罕见但关键的驾驶场景严重缺失。现有方案如安全关键场景生成和闭环学习多依赖规则启发式、重采样或离线数据训练的生成模型,难以产生多样且新颖的挑战。近期工作虽利用视觉语言模型(VLM)生成场景描述以指导下游模型生成危险轨迹,但两阶段框架限制了VLM的生成潜力,最终轨迹多样性受限于下游算法的泛化上限。为此,我们提出VILTA(VLM-In-the-Loop Trajectory Adversary),一种将VLM嵌入自动驾驶代理闭环训练的新框架。不同于以往方法,VILTA通过理解动态驾驶环境,直接精细编辑周围代理的未来轨迹,主动参与训练过程。该直接编辑机制充分发挥VLM的强大泛化能力,生成多样化且合理的挑战性场景,超越传统方法范围。实验表明,该方法显著提升所获自动驾驶策略的安全性与鲁棒性,尤其在应对关键长尾事件方面表现突出。

原文摘要 · Abstract (English)

The safe deployment of autonomous driving (AD) systems is fundamentally hindered by the long-tail problem, where rare yet critical driving scenarios are severely underrepresented in real-world data. Existing solutions including safety-critical scenario generation and closed-loop learning often rely on rule-based heuristics, resampling methods and generative models learned from offline datasets, limiting their ability to produce diverse and novel challenges. While recent works leverage Vision Language Models (VLMs) to produce scene descriptions that guide a separate, downstream model in generating hazardous trajectories for agents, such two-stage framework constrains the generative potential of VLMs, as the diversity of the final trajectories is ultimately limited by the generalization ceiling of the downstream algorithm. To overcome these limitations, we introduce VILTA (VLM-In-the-Loop Trajectory Adversary), a novel framework that integrates a VLM into the closed-loop training of AD agents. Unlike prior works, VILTA actively participates in the training loop by comprehending the dynamic driving environment and strategically generating challenging scenarios through direct, fine-grained editing of surrounding agents' future trajectories. This direct-editing approach fully leverages the VLM's powerful generalization capabilities to create a diverse curriculum of plausible yet challenging scenarios that extend beyond the scope of traditional methods. We demonstrate that our approach substantially enhances the safety and robustness of the resulting AD policy, particularly in its ability to navigate critical long-tail events.

自动驾驶视觉语言模型对抗生成长尾问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。