arXiv:2602.23898cs.CVcs.AI2026-02被引 10

新基准测试揭露多模态模型在指代表达理解中的推理短板

Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks

  • 设计含复杂语言表达和强干扰项的指代表达数据集
  • 模型在新任务上性能大幅下降,暴露对捷径的依赖
  • 适合关注视觉推理与真实理解的AI研究者使用

指代表达理解(REC)将语言与视觉区域感知关联。现有基准(RefCOCO、RefCOCO+、RefCOCOg)虽因多模态大模型取得进展,但难以有效检验视觉推理与定位能力:(i)多数表达过短,推理需求低;(ii)图像干扰项少,目标易定位;(iii)冗余描述允许模型绕过真正理解。本文提出Ref-Adv,通过仅保留唯一确定目标所需信息的非平凡语言表达,抑制捷径。数据集包含真实图像上的指代表达,配有高难度干扰项及包含否定等推理维度的标注。通过词序扰动与描述符删除实验,证明解决Ref-Adv需超越简单线索的推理。对多种主流多模态大模型评估发现,尽管在旧基准表现良好,但在Ref-Adv上性能显著下降,揭示其对捷径的依赖及视觉推理与定位能力的不足。提供深入失败分析,旨在引导未来多模态模型在视觉推理与定位方面的研究。

原文摘要 · Abstract (English)

Referring Expression Comprehension (REC) links language to region level visual perception. Standard benchmarks (RefCOCO, RefCOCO+, RefCOCOg) have progressed rapidly with multimodal LLMs but remain weak tests of visual reasoning and grounding: (i) many expressions are very short, leaving little reasoning demand; (ii) images often contain few distractors, making the target easy to find; and (iii) redundant descriptors enable shortcut solutions that bypass genuine text understanding and visual reasoning. We introduce Ref-Adv, a modern REC benchmark that suppresses shortcuts by pairing linguistically nontrivial expressions with only the information necessary to uniquely identify the target. The dataset contains referring expressions on real images, curated with hard distractors and annotated with reasoning facets including negation. We conduct comprehensive ablations (word order perturbations and descriptor deletion sufficiency) to show that solving Ref-Adv requires reasoning beyond simple cues, and we evaluate a broad suite of contemporary multimodal LLMs on Ref-Adv. Despite strong results on RefCOCO, RefCOCO+, and RefCOCOg, models drop markedly on Ref-Adv, revealing reliance on shortcuts and gaps in visual reasoning and grounding. We provide an in depth failure analysis and aim for Ref-Adv to guide future work on visual reasoning and grounding in MLLMs.

视觉推理多模态指代表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。