arXiv:2606.24335cs.CV2026-06被引 1

用单图测物体尺寸,暴露视觉语言模型的证据依赖漏洞。

Ill-Posed by Design: Probing Evidence Use in VLMs

论文配图:Ill-Posed by Design: Probing Evidence Use in VLMs
图 1 · 摘自论文原文
  • 设计单图尺寸估计任务,让模型依赖模糊线索判断真实大小。
  • 大模型在野外场景仍落后于纯文本模型,最大仅达92%准确率。
  • 模型严重依赖目标类别和像素,却忽略场景几何结构。

反事实分析广泛用于研究视觉语言模型的证据使用,但在良构任务中诊断价值有限:当多个线索独立支持同一答案时,移除其中一个可能不影响预测。我们提出单目度量物体尺寸估计作为非良构诊断设置:因单张未校准图像无法确定真实尺寸,模型必须依赖不完善的线索——类别先验、目标外观、局部上下文、表观尺寸和场景几何。我们构建了包含10,813个维度查询(来自Objectron)和331个实地测量场景的度量VQA数据集,并评估12个开源权重的VLM(参数量3–397B)。即使最大的模型(Qwen3-VL-235B、Qwen3.5-397B、InternVL3.5-241B)在野外数据集上仍落后于纯文本前沿模型。反事实分析显示:目标身份是最重要的线索;目标像素与局部上下文仅对部分模型有帮助;表观尺寸可改变预测但无方向性输出;全局场景几何几乎未被利用。我们进一步分析LoRA微调作为针对度量估计的具体干预手段:虽然任务可学习,但模型未能学会利用场景几何。

原文摘要 · Abstract (English)

Counterfactual analysis is widely used to study evidence use in vision-language models, but its diagnostic value is limited on well-posed tasks: when several cues independently support the same answer, removing one may not change the prediction. We propose monocular metric object-size estimation as an ill-posed diagnostic setting for evidence selection: because physical size cannot be determined from a single uncalibrated image, models must rely on imperfect cues category priors, target appearance, local context, apparent image size, and scene geometry. We assemble Metric VQA ($10{,}813$ dimension queries from Objectron and $331$ tape-measured in-the-wild scenes) and evaluate $12$ open-weight VLMs ($3$--$397$\,B parameters) with counterfactual analysis decomposing six visual and language evidence channels. Even the largest VLMs tested (Qwen3-VL-235B, Qwen3.5-397B, InternVL3.5-241B) trail a text-only frontier LLM on the in-the-wild split. The diagnostic analysis shows: target identity is the most load-bearing cue, target pixels and local context help only some models, apparent size shifts predictions without a directional readout, and global scene geometry is largely unused. We analyze LoRA fine-tuning as an actionable intervention specific to metric estimation: while the task is learnable, the models do not learn to leverage scene geometry.

视觉语言模型证据推理度量估计反事实分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。