arXiv:2509.01209cs.CV2025-09ICCV

提出无参考评估方法,提升开放词汇场景图生成模型的评测公平性。

Measuring Image-Relation Alignment: Reference-Free Evaluation of VLMs and Synthetic Pre-training for Open-Vocabulary Scene Graph Generation

  • 设计无需参考的评估指标,公平衡量VLM在开放词汇关系预测中的能力。
  • 用区域特异性提示调优生成高质量合成数据,提升预训练效果。
  • 适合研究开放词汇视觉语言模型与场景图生成的开发者和研究人员。

场景图生成(SGG)将图像中物体间的关系编码为图结构。得益于视觉-语言模型(VLMs)的发展,开放词汇场景图生成任务应运而生,旨在评估模型对广泛多样关系的学习能力。然而,现有基准词汇量有限,导致对开源模型的评估效率低下。本文提出一种新的无参考评估指标,可公平评估VLM在开放词汇关系预测中的表现。此外,开放词汇SGG还依赖于质量较差的弱监督数据进行预训练。为此,我们提出通过区域特异性提示调优快速生成高质量合成数据的新方法。实验表明,使用该合成数据进行预训练能显著提升开放词汇SGG模型的泛化能力。

原文摘要 · Abstract (English)

Scene Graph Generation (SGG) encodes visual relationships between objects in images as graph structures. Thanks to the advances of Vision-Language Models (VLMs), the task of Open-Vocabulary SGG has been recently proposed where models are evaluated on their functionality to learn a wide and diverse range of relations. Current benchmarks in SGG, however, possess a very limited vocabulary, making the evaluation of open-source models inefficient. In this paper, we propose a new reference-free metric to fairly evaluate the open-vocabulary capabilities of VLMs for relation prediction. Another limitation of Open-Vocabulary SGG is the reliance on weakly supervised data of poor quality for pre-training. We also propose a new solution for quickly generating high-quality synthetic data through region-specific prompt tuning of VLMs. Experimental results show that pre-training with this new data split can benefit the generalization capabilities of Open-Voc SGG models.

场景图生成视觉语言模型开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。