arXiv:2512.17495cs.CV2025-12被引 10

新基准揭示大模型视觉定位能力严重不足,尤其无法识别无法定位的提问。

GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation

  • 从区分相似物体、空间关系、遮挡对象到拒绝不合理问题,四维测试模型真实能力。
  • 顶尖模型准确率仅45.1%,多数模型在拒绝任务上为0%。
  • 适合关注多模态模型真实理解力的研究者与开发者参考。

视觉定位是将自然语言描述中的对象定位到图像中的关键任务,是连接语言与视觉理解的桥梁。尽管多模态大语言模型(MLLMs)在现有基准上表现优异,但其是否具备类人水平的视觉定位能力仍存疑。当前基准难以反映真实世界的复杂性,而人类能轻松处理复杂指代并识别无法定位的情况。为此,我们提出GroundingME,一个系统性评估模型在四个维度上的基准:(1) 区分性:辨别高度相似物体;(2) 空间性:理解复杂空间关系描述;(3) 限制性:处理遮挡或微小物体;(4) 拒绝性:识别不可定位查询。通过自动化生成结合人工验证,构建了1,005个高挑战性样本,贴近真实场景。评估25个主流MLLM发现显著能力差距:最优模型准确率为45.1%,多数模型在拒绝任务上得分为0%。我们探索两种改进策略:(1) 测试时缩放通过选择最佳推理轨迹,提升整体性能最多4.5%;(2) 数据混合训练使拒绝准确率从0%提升至27.9%。GroundingME既是诊断工具,也指明通往人类级视觉定位的路径。

原文摘要 · Abstract (English)

Visual grounding, localizing objects from natural language descriptions, represents a critical bridge between language and vision understanding. While multimodal large language models (MLLMs) achieve impressive scores on existing benchmarks, a fundamental question remains: can MLLMs truly visually ground with human-like sophistication, or are they merely pattern-matching on simplified datasets? Current benchmarks fail to capture real-world complexity where humans effortlessly navigate intricate references and recognize when grounding is impossible. To rigorously assess MLLMs' true capabilities, we introduce GroundingME, a benchmark that systematically challenges models across four critical dimensions: (1) Discriminative: distinguishing highly similar objects, (2) Spatial: understanding complex relational descriptions, (3) Limited: handling occlusions or tiny objects, and (4) Rejection: recognizing ungroundable queries. Through careful curation combining automated generation with human verification, we create 1,005 challenging examples mirroring real-world complexity. Evaluating 25 state-of-the-art MLLMs reveals a profound capability gap: the best model achieves only 45.1% accuracy, while most score 0% on rejection tasks. We explore two strategies for improvements: (1) test-time scaling selects optimal response by thinking trajectory to improve overall performance by up to 4.5%, and (2) data-mixture training boosts rejection accuracy from 0% to 27.9%. GroundingME thus serves as both a diagnostic tool revealing current limitations in MLLMs and a roadmap toward human-level visual grounding. Project page: https://groundingme.github.io

视觉定位多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。