arXiv:2604.03765cs.CV2026-04

提出新框架ITIScore,自动评估多模态模型图文生成质量。

ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs

论文配图:ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs
图 1 · 摘自论文原文
  • 用图像-文本-图像重建一致性衡量字幕质量
  • 4万条跨12类图像的字幕,含长短句与人类评分
  • 可零样本适配其他数据集,适合评测图文模型

近年来多模态大语言模型在图像理解与字幕生成方面取得显著进展。然而现有字幕评估基准普遍存在字幕长度单一、缺少最新先进模型、人工标注不足等问题,可能导致偏差并限制对现代多模态模型性能的全面评估。为此,我们提出一个新的大规模字幕基准ICBench,涵盖12个内容类别,在2000张图像上由10个先进多模态大模型生成长短字幕,共4万条。我们开展广泛的人类主观评价,获取细粒度维度的平均意见分数(MOS),短字幕评估流畅性、相关性与简洁性,长字幕则基于流畅性、相关性与完整性。此外,我们提出一种基于图像-文本-图像框架的自动化评估指标ITIScore,通过重建一致性衡量字幕质量。实验表明该指标与人工判断高度一致,并在其他公开字幕数据集上展现出强零样本泛化能力。数据集与模型将在发表后开源。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing image captioning benchmarks typically suffer from limited diversity in caption length, the absence of recent advanced MLLMs, and insufficient human annotations, which potentially introduces bias and limits the ability to comprehensively assess the performance of modern MLLMs. To address these limitations, we present a new large-scale image captioning benchmark, termed, ICBench, which covers 12 content categories and consists of both short and long captions generated by 10 advanced MLLMs on 2K images, resulting in 40K captions in total. We conduct extensive human subjective studies to obtain mean opinion scores (MOSs) across fine-grained evaluation dimensions, where short captions are assessed in terms of fluency, relevance, and conciseness, while long captions are evaluated based on fluency, relevance, and completeness. Furthermore, we propose an automated evaluation metric, \textbf{ITIScore}, based on an image-to-text-to-image framework, which measures caption quality through reconstruction consistency. Experimental results demonstrate strong alignment between our automatic metric and human judgments, as well as robust zero-shot generalization ability on other public captioning datasets. Both the dataset and model will be released upon publication.

图文生成评估框架多模态自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。