让模型理解场景中角色与意图,而非仅匹配名词。
Beyond Referring Expressions: Scenario Comprehension Visual Grounding
- 用段落级描述评估角色、目标和上下文关系的视觉定位能力。
- 31k训练+4k测试数据,包含未见类别和干扰项,揭示模型深层缺陷。
- 提出分阶段推理方法,显著提升复杂场景下的定位准确率。
现有视觉定位基准主要评估图像区域与字面指代表达的对齐,模型常通过匹配明显命名类别即可成功。本文探索更复杂且互补的场景理解型视觉定位任务,目标需从角色、意图和关系上下文中推断,而非依赖显式命名。为此提出参考场景理解(RSC)基准,其查询为描述对象角色、用户目标和上下文线索的段落文本,包含故意设置的干扰对象,需深度理解才能解析。每个样本标注了可解释的难度标签(唯一性、杂乱度、大小、重叠、位置),暴露不同失败模式并支持细粒度分析。RSC包含约31,000个训练样本、4,000个域内测试样本及3,000个分布外样本(含未见物体类别)。进一步提出ScenGround,一种课程推理方法,结合监督预热与难度感知强化学习。实验表明,场景型查询暴露出当前模型在标准基准中未揭示的系统性缺陷,而课程训练在困难子集上提升性能,并可迁移至标准基准。
原文摘要 · Abstract (English)
Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent named category. We explore a complementary and more challenging setting of scenario-based visual grounding, where the target must be inferred from roles, intentions, and relational context rather than explicit naming. We introduce Referring Scenario Comprehension (RSC), a benchmark designed for this setting. The queries in this benchmark are paragraph-length texts that describe object roles, user goals, and contextual cues, including deliberate references to distractor objects that often require deep understanding to resolve. Each instance is annotated with interpretable difficulty tags for uniqueness, clutter, size, overlap, and position which expose distinct failure modes and support fine-grained analysis. RSC contains approximately 31k training examples, 4k in-domain test examples, and a 3k out-of-distribution split with unseen object categories. We further propose ScenGround, a curriculum reasoning method serving as a reference point for this setting, combining supervised warm-starting with difficulty-aware reinforcement learning. Experiments show that scenario-based queries expose systematic failures in current models that standard benchmarks do not reveal, and that curriculum training improves performance on challenging slices and transfers to standard benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。