arXiv:2503.21721cs.CV2025-03被引 1

提出统一评估文生图/视频质量与文本一致性的新指标cFreD。

Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance

  • 基于条件弗雷歇距离,整合视觉保真度与文本一致性评估。
  • 在多个模型和提示数据集上,与人类判断相关性高于现有方法。
  • 无需重新训练,适用于新模型和分布外提示,适合评测人员使用。

文生图和文生视频模型的评估面临根本性挑战:现有指标无法同时衡量视觉质量和与文本提示的一致性,导致与人类判断的相关性较差。为此,我们提出cFreD,一种基于条件弗雷歇距离的通用评估指标,将视觉保真度与文本条件一致性统一为单一评分。传统指标如FID仅关注图像质量而忽略文本条件,而CLIPScore等对视觉质量不敏感。此外,基于学习偏好模型需持续重训练,难以泛化至新架构或分布外提示。在多个近期提出的文生图模型及多样化提示数据集上的大量实验表明,cFreD相较于统计指标(包括基于人类偏好的指标)与人类判断具有更高相关性。结果验证了cFreD作为稳健、可未来扩展的评估工具,能标准化该快速演进领域的基准测试。我们已发布评估工具包与基准数据集。

原文摘要 · Abstract (English)

Evaluating text-to-image and text-to-video models is challenging due to a fundamental disconnect: established metrics fail to jointly measure visual quality and semantic alignment with text, leading to a poor correlation with human judgments. To address this critical issue, we propose cFreD, a general metric based on a Conditional Fréchet Distance that unifies the assessment of visual fidelity and text-prompt consistency into a single score. Existing metrics such as Fréchet Inception Distance (FID) capture image quality but ignore text conditioning while alignment scores such as CLIPScore are insensitive to visual quality. Furthermore, learned preference models require constant retraining and are unlikely to generalize to novel architectures or out-of-distribution prompts. Through extensive experiments across multiple recently proposed text-to-image models and diverse prompt datasets, cFreD exhibits a higher correlation with human judgments compared to statistical metrics , including metrics trained with human preferences. Our findings validate cFreD as a robust, future-proof metric for the systematic evaluation of text conditioned models, standardizing benchmarking in this rapidly evolving field. We release our evaluation toolkit and benchmark.

生成评估文本生成图像生成指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。