arXiv:2608.12088cs.CV2026-08

新评估方法让生成图像质量分析更直观,能看懂位置和属性是否对得上。

RA-ClipScore: Making Generative Model Evaluation More Interpretable

论文配图:RA-ClipScore: Making Generative Model Evaluation More Interpretable
图 1 · 摘自论文原文
  • 用双提示词解耦属性,通过局部区域特征分析图像空间分布
  • 在分布错位和无关文本条件下仍保持评估可靠性,比旧方法更准
  • 适合研究生成模型偏差、做视觉多样性评估的开发者与评测者

生成模型可产出几乎无法与真实数据区分的图像,但严谨且可解释的评估仍具挑战。传统指标如FID仅提供单一分数,缺乏诊断信息。基于CLIP的指标虽能进行语义评估,但受限于CLIP训练范式,难以实现属性层面的细粒度分析。我们提出RA-CLIPScore,该方法缓解上述问题,将基于CLIP的评估扩展至空间分布对齐,衡量生成对象是否符合训练数据中的位置先验。RA-CLIPScore引入双提示词以解耦竞争性属性,并利用局部图像块特征捕捉细粒度区域语义。我们在生成模型匹配训练数据属性与空间分布的能力上进行了评估。大量实验表明,相较于已有方法,RA-CLIPScore提供了更稳健、更可解释的评估结果,尤其在分布错位或部分无关文本属性情况下表现优异。我们进一步展示了其揭示生成模型空间偏见的能力。用户评估证实,基于RA-CLIPScore的区域单属性差异(Regional Single Attribute Divergence)与人类对视觉多样性的感知更一致。

原文摘要 · Abstract (English)

Generative models can produce images nearly indistinguishable from real data, yet rigorous and interpretable evaluation remains challenging. Conventional metrics such as FID provide only scalar scores with limited diagnostic insight. Widely adopted CLIP-based metrics enable semantic evaluation beyond simple training class labels, but inherit limitations from CLIP's training paradigm that restrict attribute-wise analysis. We propose RA-CLIPScore, a novel metric that mitigates these issues and extends CLIP-based evaluation to spatial distribution alignment, measuring whether generated objects adhere to the positional priors found in the training data. RA-CLIPScore introduces dual prompts to decouple competing attributes and leverages local patch tokens to capture fine-grained regional semantics. We evaluate image generative models on their ability to match both attribute and spatial distributions of the training data. Extensive experiments show that RA-CLIPScore provides more robust and interpretable evaluations than prior methods, particularly under distribution misalignment or partially irrelevant textual attributes. We further demonstrate how it reveals spatial biases in generative models. User evaluations confirm that Regional Single Attribute Divergence based on our RA-CLIPScore aligns more closely with human perception of visual diversity than existing semantic metrics.

图像评估生成模型可解释性空间分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。