arXiv:2511.19914cs.RO2025-11NeurIPS被引 5

用对抗迁移让仿真数据教会真实驾驶系统处理罕见场景

CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model

  • 构建因果链视觉语言模型,端到端推理复杂驾驶逻辑
  • 通过对抗训练使学生模型从仿真数据学会处理长尾场景
  • 适合关注自动驾驶可解释性与泛化能力的研究者

自动驾驶是人工智能的重要应用。当前方法已从常规场景转向处理复杂、长尾的罕见情况,如细微人类行为、交通事故和违规驾驶。得益于大语言模型在视觉与自然语言理解及指令遵循方面的表现,近期研究将大模型融入自动驾驶系统以提升推理能力、可解释性与跨场景性能。然而现有方法通常依赖真实世界数据(适合工业部署)或针对罕见场景定制的仿真数据,缺乏有效融合两者优势的机制。为此,本文提出一种新型端到端对抗迁移框架CoC-VLA,实现从仿真环境向真实世界长尾场景处理能力的迁移。该框架包含教师视觉语言模型、学生视觉语言模型和判别器。教师与学生模型共享基于因果链视觉语言模型(CoC VLM)的基底架构,通过端到端文本适配器整合时序信息,支持链式思维推理以推断复杂驾驶逻辑。教师模型在仿真数据上预训练,学生模型在真实数据上预训练。判别器通过新型反向传播策略,促使学生模型从教师模型中学习长尾场景应对能力,完成跨域知识迁移。

原文摘要 · Abstract (English)

Autonomous driving represents a prominent application of artificial intelligence. Recent approaches have shifted from focusing solely on common scenarios to addressing complex, long-tail situations such as subtle human behaviors, traffic accidents, and non-compliant driving patterns. Given the demonstrated capabilities of large language models (LLMs) in understanding visual and natural language inputs and following instructions, recent methods have integrated LLMs into autonomous driving systems to enhance reasoning, interpretability, and performance across diverse scenarios. However, existing methods typically rely either on real-world data, which is suitable for industrial deployment, or on simulation data tailored to rare or hard case scenarios. Few approaches effectively integrate the complementary advantages of both data sources. To address this limitation, we propose a novel VLM-guided, end-to-end adversarial transfer framework for autonomous driving that transfers long-tail handling capabilities from simulation to real-world deployment, named CoC-VLA. The framework comprises a teacher VLM model, a student VLM model, and a discriminator. Both the teacher and student VLM models utilize a shared base architecture, termed the Chain-of-Causality Visual-Language Model (CoC VLM), which integrates temporal information via an end-to-end text adapter. This architecture supports chain-of-thought reasoning to infer complex driving logic. The teacher and student VLM models are pre-trained separately on simulated and real-world datasets. The discriminator is trained adversarially to facilitate the transfer of long-tail handling capabilities from simulated to real-world environments by the student VLM model, using a novel backpropagation strategy.

自动驾驶视觉语言模型对抗迁移可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。