arXiv:2510.19060cs.CVcs.AI2025-10中稿 · ICLR被引 1

用场景图引导大模型评分,让图像描述评估更精准。

PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions

  • 以场景图为结构化标准,指导大模型进行细粒度错误判断。
  • 在艺术图像数据集上与人工评分相关性提升0.05(Spearman ρ)。
  • 适合评估复杂图像描述任务,尤其对视觉语言模型性能测试有帮助。

尽管视觉语言模型在生成详细图像描述方面取得进展,但评估仍具挑战。传统指标(如CIDEr、SPICE)针对短文本设计,对如今已少见的物体误识等错误敏感,而长文本需关注属性与关系的准确性,并定位具体错误段落。本文提出PoSh,一种利用场景图作为结构化评分标准,引导大模型作为裁判的评估指标,可生成基于细粒度错误(如组合理解错误)的聚合得分。PoSh具备可复现性与可解释性,相比现有指标(包括GPT4o-as-a-Judge)更贴近人工评分。为验证其有效性,我们构建了新基准DOCENT,包含艺术作品及其专家参考描述,以及学生给出的细粒度与粗粒度质量评价。实验表明,PoSh在DOCENT中与人类评分的相关性比最佳开源替代方案高0.05(Spearman ρ),在不同图像类型下表现稳健(使用CapArena数据集验证),且可作为有效奖励函数,优于标准监督微调。使用PoSh分析发现,基础模型在描述含丰富场景动态的绘画、素描和雕像时仍存在系统性遗漏,揭示了一个新的高难度评测任务,以推动视觉语言模型发展。

原文摘要 · Abstract (English)

While vision-language models (VLMs) have advanced into detailed image description, evaluation remains a challenge. Standard metrics (e.g. CIDEr, SPICE) were designed for short texts and tuned to recognize errors that are now uncommon, such as object misidentification. In contrast, long texts require sensitivity to attribute and relation attachments and scores that localize errors to particular text spans. In this work, we introduce PoSh, a metric for detailed image description that uses scene graphs as structured rubrics to guide LLMs-as-a-Judge, producing aggregate scores grounded in fine-grained errors (e.g. mistakes in compositional understanding). PoSh is replicable, interpretable and a better proxy for human raters than existing metrics (including GPT4o-as-a-Judge). To validate PoSh, we introduce a challenging new dataset, DOCENT. This novel benchmark contains artwork, paired with expert-written references, and model-generated descriptions, augmented with granular and coarse judgments of their quality from art history students. Thus, DOCENT enables evaluating both detailed image description metrics and detailed image description itself in a challenging new domain. We show that PoSh achieves stronger correlations (+0.05 Spearman $ρ$) with the human judgments in DOCENT than the best open-weight alternatives, is robust to image type (using CapArena, an existing dataset of web imagery) and is a capable reward function, outperforming standard supervised fine-tuning. Then, using PoSh, we characterize the performance of open and closed models in describing the paintings, sketches and statues in DOCENT and find that foundation models struggle to achieve full, error-free coverage of images with rich scene dynamics, establishing a demanding new task to gauge VLM progress. Through both PoSh and DOCENT, we hope to enable advances in important areas such as assistive text generation.

图像描述大模型评估场景图多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。