首次系统检验自动驾驶视觉语言模型推理可靠性,发现其逻辑与现实严重不符。
Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models
- 通过300次推断分析100个场景,用信息论定义推理忠实性
- 仅42.5%推理与真实因果链匹配,94名行人被漏检
- 模型对微小干扰极度敏感,近半数决策逻辑不一致
本文首次系统研究视觉-语言-动作(VLA)驾驶模型的推理忠实性,分析了在100个多样化的PhysicalAI-AV场景中300次Alpamayo-R1-10B的推断结果。主要发现显示:输出的自然语言推理与轨迹可能严重失真——(i)整体推理忠实度仅为42.5%,因果链与现实匹配不足一半;(ii)在三分之一涉及行人的场景中漏检94名行人;(iii)在轻微视觉扰动下,97.7%的轨迹表现出脆弱性;(iv)推理与行为一致性均值仅48.3%,其中53.3%的推断存在低一致性,包括37.9%声称要刹车却继续前行的情况。本文从信息论角度形式化忠实性,定义实体与动作忠实性验证标准,并提出一套四组件安全架构以应对上述问题。
原文摘要 · Abstract (English)
We present the first systematic study of faithfulness in Vision-Language-Action (VLA) driving models, analyzing 300 Alpamayo-R1-10B inferences across 100 diverse PhysicalAI-AV scenarios. Our main finding is that output natural-language rationales with trajectories may be significantly unfaithful: (i) overall reasoning fidelity is only 42.5%, with Chain-of-Causation matching scene reality less than half the time; (ii) 94 missed pedestrians in one-third of pedestrian-relevant scenes; (iii) 97.7% trajectory fragility under mild visual perturbations; and (iv) only 48.3% mean reasoning-action consistency, with 53.3% of inferences exhibiting low consistency, including 37.9% of stop-claimed cases where the model continues instead. We formalize faithfulness information-theoretically, define entity and action fidelity with verification criteria, and outline a four-component safety architecture aligned with these results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。