现有图文评价指标对语义不变性不敏感,微小修改却大幅改变评分。
Do Image-Text Metrics Respect Semantic Invariances?

- 通过空间、物体和语义框架三类扰动测试五种主流评价指标
- 平均评分波动6%~9%,0.7%差异时37%情况导致排名翻转
- 提出后处理校准方法,使敏感度减半且保持与专家评分相关性
无参考的图文对齐评价指标已成为标准,但其是否尊重语义不变性尚不明确。我们针对五种流行评估器(CLIPScore、PAC-S、UMIC、FLEUR及确定性LLM裁判)在三个维度的语义保持扰动下进行不变性探测:空间(翻转、上下文保持重置、轻度旋转)、物体(尺度、类别)以及社会语言框架(文化/经济形容词,配以中性且长度匹配的对照)。在三个检测数据集和三个标题评价套件的精选子集上,发现一致的非语义敏感性:轻微的空间修改和简单措辞变化平均导致评分波动6%~9%,系统间仅相差0.7%时,多达约37%情况下引发排名翻转,尤其在空间扰动下更显著。小规模人类实验支持该发现,且标注者普遍认为扰动对等正确,说明评分波动反映的是度量行为而非语义变化。我们进一步提出不变性校准评分,一种后处理调整方法,可将中位绝对敏感度大致减半,同时保留与学习型标题评价器的相关性。
原文摘要 · Abstract (English)
Reference-free image-to-text evaluators are now standard for scoring image-caption alignment, yet it is unclear whether they respect semantic invariances. We present an invariance probe on five popular evaluators (CLIPScore, PAC-S, UMIC, FLEUR, and a deterministic LLM judge) under semantics-preserving perturbations along three axes -- spatial (flips, context-preserving repositioning, light rotations), object (scale, category), and socio-linguistic framing (cultural/economic adjectives with neutral and length-matched controls). Across curated slices of three detection datasets and three caption evaluation suites, we find consistent non-semantic sensitivities, where benign spatial edits and simple phrasing changes shift scores by $\approx$6--9\% on average, and for systems separated by just 0.7\%, these shifts can cause ranking flips in up to $\sim$37\% of cases, particularly under spatial changes. A small human study also supports this finding and confirms that annotators generally judge perturbed pairs as equally correct, so these shifts reflect metric behavior rather than semantic change. We further propose invariance-calibrated scoring, a post-hoc adjustment that roughly halves median absolute sensitivity while retaining correlation with learned caption evaluators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。