构建图文理解新基准,用有向场景图评估图像描述的全面性
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
- 基于有向场景图构建细粒度图像描述评测框架
- 在CompreCap数据集上验证了评估结果与人工评分高度一致
- 适合研究视觉语言模型细节理解能力的学者使用
生成包含丰富文本信息的详细图像描述已成为大视觉-语言模型(LVLMs)的重要研究方向。然而,现有研究缺乏专门用于评估详细描述准确性和全面性的基准。本文提出一个名为CompreCap的详细描述评测基准,从有向场景图视角评估视觉上下文理解能力。具体而言,首先根据常见物体词汇对图像进行语义区域分割(即语义分割掩码),并区分各区域内物体的属性;随后标注物体间的有向关系,构成能充分编码图像复合信息的有向场景图。基于该场景图,我们设计了一套多层级评估流程,涵盖对象覆盖度、属性描述准确性、关键关系得分等维度。在CompreCap数据集上的实验结果表明,本评估方法与人工评价分数在不同LVLM间高度一致。
原文摘要 · Abstract (English)
Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed captions to measure their accuracy and comprehensiveness. In this paper, we introduce a detailed caption benchmark, termed as CompreCap, to evaluate the visual context from a directed scene graph view. Concretely, we first manually segment the image into semantically meaningful regions (i.e., semantic segmentation mask) according to common-object vocabulary, while also distinguishing attributes of objects within all those regions. Then directional relation labels of these objects are annotated to compose a directed scene graph that can well encode rich compositional information of the image. Based on our directed scene graph, we develop a pipeline to assess the generated detailed captions from LVLMs on multiple levels, including the object-level coverage, the accuracy of attribute descriptions, the score of key relationships, etc. Experimental results on the CompreCap dataset confirm that our evaluation method aligns closely with human evaluation scores across LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。