arXiv:2605.06080cs.CV2026-05

无需参考文本,通过多尺度分布匹配评估图像描述质量

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation

论文配图:MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
图 1 · 摘自论文原文
  • 将图像与文本嵌入建模为球面混合分布,捕捉细粒度语义差异
  • 在多个候选描述上实现领先的人类判断相关性,优于现有无参考指标
  • 可分解诊断局部错误,适合需要透明评估的场景

无参考图像描述评估仍具挑战性,因全局嵌入相似性常忽略细粒度不匹配,如幻觉物体、缺失属性或关系错误。本文提出MSD-Score,一种无参考度量,将图像块和文本标记嵌入建模为单位超球面上的冯·米塞斯-费舍尔混合分布。不同于将每种模态视为单一数据点,MSD-Score将图像-文本匹配形式化为多尺度分布评分问题。通过加权双向KL散度量化语义差异,并在多尺度框架中融合全局相似性,支持单候选与多候选评估。大量实验表明,MSD-Score在无参考指标中实现了与人类判断的最佳相关性。除精度外,其概率化框架提供透明且可分解的局部定位错误诊断,为整体相似性度量和人工评判提供确定性补充信号。

原文摘要 · Abstract (English)

Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attributes, or incorrect relations. We propose MSD-Score, a reference-free metric that models image patch and text token embeddings as von Mises-Fisher mixtures on the unit hypersphere. Instead of treating each modality as a single point, MSD-Score formulates image-text matching as a multi-scale distributional scoring problem. Semantic discrepancies are quantified via a weighted bi-directional KL divergence and combined with global similarity in a multi-scale framework for both single- and multi-candidate evaluations. Extensive experiments show that MSD-Score achieves state-of-the-art correlation with human judgments among reference-free metrics. Beyond accuracy, its probabilistic formulation yields transparent and decomposable diagnostics of local grounding errors, providing a deterministic complementary signal to holistic similarity metrics and judge-based evaluators.

图像描述无参考评估分布建模可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。