arXiv:2412.10604cs.CV2024-12被引 7

统一评估生成图像模型,支持多数据集与指标灵活扩展。

EvalGIM: A Library for Evaluating Generative Image Models

  • 构建可插拔的评估框架,兼容多种数据集与评价指标。
  • 提供两种前沿评估方法,实现模型性能的精准对比。
  • 内置可复现分析工具,助力发现模型真实优劣。

随着文本到图像生成模型的广泛应用,自动化基准测试方法也日益普及。然而,尽管评价指标和数据集众多,却缺乏统一的基准测试库来支持跨数据集和指标的评估。同时,新评估方法快速迭代,要求评估工具具备高度灵活性。此外,如何整合评估结果以获得可操作的性能洞察仍存在空白。为此,我们推出了EvalGIM(发音为"EvalGym"),一个用于评估生成图像模型的库。EvalGIM全面支持衡量图像质量、多样性与一致性的各类数据集和指标。其设计优先考虑用户自定义能力,支持新数据集与指标的即插即用。为实现可操作的评估洞见,我们引入了“评估任务”(Evaluation Exercises),包含两种最先进的文本到图像生成模型评估方法:一致性-多样性-真实性帕累托前沿分析,以及跨群体性能差异的解耦测量。此外,还新增了模型排名鲁棒性分析和不同提示风格下的均衡评估方法。我们鼓励使用EvalGIM探索文本到图像模型,并欢迎在https://github.com/facebookresearch/EvalGIM/贡献代码。

原文摘要 · Abstract (English)

As the use of text-to-image generative models increases, so does the adoption of automatic benchmarking methods used in their evaluation. However, while metrics and datasets abound, there are few unified benchmarking libraries that provide a framework for performing evaluations across many datasets and metrics. Furthermore, the rapid introduction of increasingly robust benchmarking methods requires that evaluation libraries remain flexible to new datasets and metrics. Finally, there remains a gap in synthesizing evaluations in order to deliver actionable takeaways about model performance. To enable unified, flexible, and actionable evaluations, we introduce EvalGIM (pronounced ''EvalGym''), a library for evaluating generative image models. EvalGIM contains broad support for datasets and metrics used to measure quality, diversity, and consistency of text-to-image generative models. In addition, EvalGIM is designed with flexibility for user customization as a top priority and contains a structure that allows plug-and-play additions of new datasets and metrics. To enable actionable evaluation insights, we introduce ''Evaluation Exercises'' that highlight takeaways for specific evaluation questions. The Evaluation Exercises contain easy-to-use and reproducible implementations of two state-of-the-art evaluation methods of text-to-image generative models: consistency-diversity-realism Pareto Fronts and disaggregated measurements of performance disparities across groups. EvalGIM also contains Evaluation Exercises that introduce two new analysis methods for text-to-image generative models: robustness analyses of model rankings and balanced evaluations across different prompt styles. We encourage text-to-image model exploration with EvalGIM and invite contributions at https://github.com/facebookresearch/EvalGIM/.

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。