构建细粒度视觉定位基准,评估模型在真实场景中的感知能力
EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models

- 设计6.6千个图像-文本-掩码三元组,覆盖三大交互阶段
- 89个模型测试显示多目标计数与部件关系理解是主要瓶颈
- 适合研究具身智能与视觉语言对齐的学者使用
尽管大型视觉语言模型(VLMs)越来越多地被用作具身智能体的感知骨干,但现有基准大多依赖问答或选择题形式,使模型可利用语言先验而非展现真正的视觉定位能力。为此,我们提出EPIC-Bench(Embodied PerceptIon BenChmark),一个面向真实具身环境的细粒度视觉接地评估基准。该基准包含6.6k个精心标注的三元组(图像、文本、掩码),涵盖23个细粒度任务,覆盖具身交互流程的三个核心阶段:目标定位、导航和操作。对超过89个主流VLMs的广泛评估表明,尽管先进推理模型展现出潜力,当前VLMs在复杂视觉-文本对齐方面普遍存在困难,尤其在多目标计数、部件-整体关系理解及可操作区域检测方面存在显著瓶颈。EPIC-Bench为推动下一代视觉驱动具身模型的发展提供了坚实基础与切实洞见。
原文摘要 · Abstract (English)
While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on question-answering or multiple-choice formats. These protocols allow models to exploit linguistic priors rather than demonstrating genuine visual grounding. To address this, we present EPIC-Bench, Embodied PerceptIon BenChmark, a fine-grained grounding benchmark designed to systematically evaluate the visual perceptual capabilities of VLMs in real-world embodied environments. Comprising 6.6k meticulously annotated tuples (Image, Text, Mask), EPIC-Bench spans 23 fine-grained tasks across three core stages of the embodied interaction pipeline: Target Localization, Navigation, and Manipulation. Extensive evaluations of over 89 leading VLMs reveal that while advanced reasoning models show promise, current VLMs universally struggle with complex visual-text alignment for physical interactions. Specifically, models exhibit critical bottlenecks in multi-target counting, part-whole relationship understanding, and affordance region detection. EPIC-Bench provides a robust foundation and actionable insights for advancing the next generation of vision-driven embodied models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。