arXiv:2512.24426cs.RO2025-12被引 19

让自动驾驶模型学会反思计划是否安全,关键时刻自动修正动作。

Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning

  • 通过模拟反事实场景,让模型在执行前自我检查并修正驾驶计划。
  • 在真实数据集上提升轨迹准确率17.6%,安全指标改善20.5%。
  • 只在复杂场景启用反思机制,实现智能省力的自适应推理。

近期的视觉-语言-动作(VLA)模型通过生成中间推理轨迹提升了端到端自动驾驶的可解释性,但大多仅描述感知与意图,很少质疑计划是否安全。本文提出反事实视觉-语言-动作模型(CF-VLA),使模型能在行动前进行自我反思与修正。该框架首先生成时间分段的元动作以总结驾驶意图,再基于元动作和视觉上下文进行反事实推理,模拟潜在结果,识别不安全行为,并输出修正后的元动作指导最终轨迹生成。为高效获得自反思能力,我们设计了滚动-筛选-标注管道,从基础(非反事实)VLA的滚动中挖掘高价值场景,并标注反事实推理轨迹用于后续训练。在大规模驾驶数据集上的实验表明,CF-VLA将轨迹准确率最高提升17.6%,安全指标提升20.5%,且具备自适应思维:仅在挑战性场景启用反事实推理。通过将推理轨迹从单次描述转变为因果性自我修正信号,CF-VLA向能‘先思后行’的自主驾驶代理迈进一步。

原文摘要 · Abstract (English)

Recent reasoning-augmented Vision-Language-Action (VLA) models have improved the interpretability of end-to-end autonomous driving by generating intermediate reasoning traces. Yet these models primarily describe what they perceive and intend to do, rarely questioning whether their planned actions are safe or appropriate. This work introduces Counterfactual VLA (CF-VLA), a self-reflective VLA framework that enables the model to reason about and revise its planned actions before execution. CF-VLA first generates time-segmented meta-actions that summarize driving intent, and then performs counterfactual reasoning conditioned on both the meta-actions and the visual context. This step simulates potential outcomes, identifies unsafe behaviors, and outputs corrected meta-actions that guide the final trajectory generation. To efficiently obtain such self-reflective capabilities, we propose a rollout-filter-label pipeline that mines high-value scenes from a base (non-counterfactual) VLA's rollouts and labels counterfactual reasoning traces for subsequent training rounds. Experiments on large-scale driving datasets show that CF-VLA improves trajectory accuracy by up to 17.6%, enhances safety metrics by 20.5%, and exhibits adaptive thinking: it only enables counterfactual reasoning in challenging scenarios. By transforming reasoning traces from one-shot descriptions to causal self-correction signals, CF-VLA takes a step toward self-reflective autonomous driving agents that learn to think before they act.

自动驾驶反事实推理自反思VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。