构建可解析的视觉语言理解评估基准,区分感知与推理能力。
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- 将指代表达分解为感知与推理两大维度,细分为六类挑战任务。
- 自动生成多样化指代表达数据,覆盖属性、位置、关系等场景。
- 提出强化学习优化方案,提升复杂推理下的定位准确率,适合模型评估者使用。
指代表达理解(REC)是一项视觉-语言任务,需根据文本描述定位图像中的特定区域。现有基准主要评估感知能力,缺乏可解释的评分机制,无法揭示多模态大模型在不同认知能力上的定位表现。为此,我们提出RefBench-PRO,一个全面的REC评估基准,将指代表达分解为感知与推理两个核心维度,并进一步细分为六个逐步增加难度的任务:属性、位置、交互、常识、关系与排除。我们还开发了全自动数据生成管道,覆盖这六类子维度,生成多样化指代表达。此外,提出Ref-R1,一种基于强化学习的训练方法,引入动态IoU引导的GRPO策略,在复杂推理条件下显著提升定位精度,建立更强基线。大量实验表明,RefBench-PRO能够可解释地评估多模态大模型在指代理解上的表现,对感知与推理均构成更大挑战。
原文摘要 · Abstract (English)
Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. Existing REC benchmarks primarily evaluate perceptual capabilities and lack interpretable scoring mechanisms, which cannot reveal the grounding capability of Multi-modal Large Language Model (MLLM) across different cognitive abilities. To address this limitation, we introduce RefBench-PRO, a comprehensive REC benchmark, which decomposes referring expressions into two core dimensions, i.e., perception and reasoning, and further subdivides them into six progressively challenging tasks, such as attribute, position, interaction, commonsense, relation and reject. We also develop a fully automated data-generation pipeline that produces diverse referring expressions across these six sub-dimensions. Furthermore, We propose Ref-R1, an RL-based learning scheme, which incorporates Dynamic IoU-based GRPO to improve localization accuracy under increasingly complex reasoning conditions, establishing a stronger baseline for REC. Extensive experiments demonstrate that our RefBench-PRO enables interpretable evaluation of MLLM on referring expression comprehension, presenting greater challenges in both perception and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。