arXiv:2608.23978cs.AIcs.CV2026-08

测试大模型在对话中逐步理解视觉目标的能力,发现其表现远低于人类。

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

论文配图:When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs
图 1 · 摘自论文原文
  • 设计多场景对话框架,模拟真实交互中的信息逐步获取过程。
  • 无初始描述时准确率最低,说明主动提问找目标仍很困难。
  • 模型自信度常高于实际表现,存在严重过拟合认知偏差。

视觉定位通常被评估为从一个明确的指代表达到视觉目标的一次性映射。这一范式忽略了现实参考的核心特征:初始指代往往不完整或模糊,需通过互动建立共同理解。本文提出一种受控的交互式视觉定位评估框架,系统调整初始信息量与需通过对话获取的信息量。在四个基于人类的视觉场景和四种交互协议下,当前大视觉语言模型(LVLMs)的表现显著低于任务级人类基线。当后续问题能修正或完善初始目标描述时,交互可提升性能;但若无初始描述,必须完全通过提问获取目标信息,性能最低,表明主动提问式定位仍具挑战。此外,模型普遍过估计自身准确性,信心与真实表现严重不符。后续研究验证了这些模式在不同描述来源(人工与AI)、推理强度、重复交互及视觉场景中的稳健性。总体而言,交互式视觉定位仍面临视觉匹配、信息搜索与综合整合的多重挑战。

原文摘要 · Abstract (English)

Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions are often incomplete or ambiguous, requiring participants to establish shared understanding through interaction. We introduce a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), varying how much target information is provided upfront and how much must be acquired through dialogue. Across four human-grounded visual contexts and four interaction protocols, current LVLMs perform significantly below task-level human baselines. Interaction can help when follow-up questions refine or repair an initial target description. Performance is lowest when no initial description is provided and target information must be acquired through questions, indicating that proactive question-driven grounding remains difficult. LVLMs are also poorly calibrated, often reporting confidence that exceeds their empirical accuracy. Follow-up studies confirm these patterns across varied description sources (human versus AI), reasoning efforts, repeated interactions, description providers, and visual contexts. Overall, interactive visual grounding remains challenging, requiring visual matching, information seeking and synthesis.

视觉定位大模型交互式评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。