新基准点出挑战:测试视觉语言模型在复杂场景中的精准定位能力
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- 构建三阶段评估体系,从指物到轨迹预测层层推进
- 十多个顶尖模型实测显示,通用大模型在精准定位上反不如部分开源模型
- 聚焦真实场景,适合研究具身智能与视觉定位的学者使用
视觉语言模型(VLMs)在多种任务中展现出强大的世界知识,被视为具身推理应用的有力候选。然而,现有基准主要通过基于图像标注的多选题评估其具身推理能力,例如选择更准确描述图像事件的路径。本文提出全新基准Point-It-Out(PIO),系统评估VLM在精确视觉定位上的具身推理能力。采用分层评估协议,涵盖三个阶段:S1(指代对象定位)、S2(任务驱动指认)、S3(视觉轨迹预测),数据来自室内、厨房、驾驶和机器人操作等具身智能关键领域。对十余个前沿VLM进行广泛实验发现:如GPT-4o等强大通用模型虽在语言、感知和推理任务表现优异,但在精确视觉定位上反而弱于某些开源模型;如MoLMO在S1、S2表现良好,但在需结合视觉轨迹规划的S3阶段表现受限。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have demonstrated impressive world knowledge across a wide range of tasks, making them promising candidates for embodied reasoning applications. However, existing benchmarks primarily evaluate the embodied reasoning ability of VLMs through multiple-choice questions based on image annotations -- for example, selecting which trajectory better describes an event in the image. In this work, we introduce the Point-It-Out (PIO) benchmark, a novel benchmark designed to systematically assess the embodied reasoning abilities of VLMs through precise visual grounding. We propose a hierarchical evaluation protocol spanning three stages (S1: referred-object localization, S2: task-driven pointing, and S3: visual trace prediction), with data collected from critical domains for embodied intelligence, including indoor, kitchen, driving, and robotic manipulation scenarios. Extensive experiments with over ten state-of-the-art VLMs reveal several interesting findings. For example, strong general-purpose models such as GPT-4o, while excelling on many benchmarks (e.g., language, perception, and reasoning), underperform compared to some open-source models in precise visual grounding; models such as MoLMO perform well in S1 and S2 but struggle in S3, where requires grounding combined with visual trace planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。