arXiv:2608.06154cs.ROcs.AI2026-08

检验视觉语言模型在零样本控制中的感知真实性,发现多数模型实际不依赖视觉输入。

Visual Grounding in Zero-Shot Vision-Language Control

论文配图:Visual Grounding in Zero-Shot Vision-Language Control
图 1 · 摘自论文原文
  • 通过输入消融实验检测模型是否真依赖视觉信息
  • 仅1个模型在镜像对称测试中达95.4%准确率,证明视觉信息可被有效利用
  • 适合关注模型可靠性与安全性的自动驾驶研究者阅读

视觉语言模型(VLMs)正被用作零样本控制器,但成功轨迹未必意味着决策真正基于视觉输入:模拟器动力学和保守动作先验可能在无真实感知的情况下产生良好得分。我们通过盲图控制、重复输入、车道轴反射、非视觉基线及管道完整性检查等输入消融实验进行验证。分析了九个直接动作模型、六个结构化局部VLMs以及一个探索性VLM-MPC层级,在两个实体平台和三个模拟器上共32,874次评分调用。直接控制结果普遍负面:恒定慢速策略优于脚本几何控制器,多个模型对图像完全不变或近乎恒定,即使能识别纵向危险的模型也无法在反射下正确变换左右指令。无一局部VLM同时满足纵向与横向接地标准。然而,仅使用图像的确定性正向控制能以0.090米平均绝对误差估计前车间隙,并实现精确镜像等价性,表明刺激物包含足够视觉信息;失败是模块化的,而非普遍现象。后处理泄漏控制对称共识守护者从16帧校准数据中选出两模型,对原始与反射视图采用2/4危险投票机制。在272个保留帧上达到0.954平衡准确率(置信区间[0.895,0.990]);嵌套留一集外验证在全部12折中复现相同模型对与阈值。弃权平局使有承诺平衡准确率达0.973,覆盖率为0.824。在保持横向控制权的前提下,离线模块重播实现0.934动作一致性与精确镜像等价性。结果支持当前VLMs作为有限、选择性的危险辅助者,而非全能零样本控制器。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception. We investigate this with an input-ablation battery: blind-image controls, repeated identical inputs, lane-axis reflection, non-visual baselines, and pipeline-integrity checks. Across nine direct-action models, six structured local VLMs, and an exploratory VLM-MPC hierarchy, we analyse 32,874 scored calls over two embodiments and three simulators. The direct-control results are largely negative: a constant-SLOW policy outperforms a scripted geometric controller, several models are image-invariant or nearly constant, and models that recognize longitudinal hazards still fail to transform LEFT and RIGHT under reflection. No local VLM meets the joint longitudinal and lateral grounding criteria. However, an image-only deterministic positive control estimates the lead gap with 0.090 m MAE and exact mirror equivariance, confirming the stimuli carry sufficient visual information; the failures are modular, not universal. A post-hoc, leakage-controlled symmetry-consensus guardian selects two models from 16 calibration frames and freezes a 2-of-4 hazard vote across original and reflected views. On 272 held-out frames it reaches 0.954 balanced accuracy (episode-cluster bootstrap 95% CI [0.895,0.990]); nested leave-one-episode-out recovers the same pair and threshold in all 12 folds. Abstaining on ties raises committed balanced accuracy to 0.973 at 0.824 coverage. With deterministic perception retaining lateral authority, offline modular replay achieves 0.934 action agreement and exact mirror equivariance. These results support current VLMs as bounded, selective hazard assistants, not monolithic zero-shot controllers.

视觉定位零样本控制模型可信度自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。