用视觉语言模型评估对象中心模型的推理能力,更真实反映其实际用途。
Evaluating Object-Centric Models beyond Object Discovery
- 用指令微调的视觉语言模型作为评估器,测试对象中心表示在复杂任务中的表现
- 提出统一指标同时衡量定位准确性和表征有用性,避免传统评估的割裂问题
- 提供多特征重建基线,帮助判断模型性能是否合理
对象中心学习(OCL)旨在学习支持组合泛化和对分布外数据鲁棒性的结构化场景表征。然而,现有方法常仅通过对象发现和简单推理任务(如图像分类探测)评估,未能充分检验其核心目标。我们指出现有基准存在两大缺陷:(1)对表征实用性缺乏足够洞察;(2)定位与表征有用性使用不一致的评估指标。为解决(1),我们采用指令微调的视觉语言模型作为评估器,在多样化的视觉问答数据集上实现可扩展的基准测试,衡量视觉语言模型如何利用OCL表征进行复杂推理。为解决(2),我们引入统一的评估任务与指标,联合评估定位(哪里)与表征有用性(是什么),消除因指标分离带来的不一致性。最后,我们引入一个简单的多特征重建基线作为参考点。
原文摘要 · Abstract (English)
Object-centric learning (OCL) aims to learn structured scene representations that support compositional generalization and robustness to out-of-distribution (OOD) data. However, OCL models are often not evaluated regarding these goals. Instead, most prior work focuses on evaluating OCL models solely through object discovery and simple reasoning tasks, such as probing the representation via image classification. We identify two limitations in existing benchmarks: (1) They provide limited insights on the representation usefulness of OCL models, and (2) localization and representation usefulness are assessed using disjoint metrics. To address (1), we use instruction-tuned VLMs as evaluators, enabling scalable benchmarking across diverse VQA datasets to measure how well VLMs leverage OCL representations for complex reasoning tasks. To address (2), we introduce a unified evaluation task and metric that jointly assess localization (where) and representation usefulness (what), thereby eliminating inconsistencies introduced by disjoint evaluation. Finally, we include a simple multi-feature reconstruction baseline as a reference point.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。