arXiv:2503.14482cs.CV2025-03ICCV被引 27

构建首个覆盖生成与编辑全链路的图像评估基准

ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing

  • 按有无源图/参考图拆解出31项细粒度任务
  • 6维11指标评估,创新引入大模型辅助编辑评测
  • 融合真实与虚拟数据,降低评估偏差,适合模型开发者使用

近年来图像生成技术进展迅速,但模型评估仍面临挑战。本文提出ICE-Bench,一个统一且全面的图像生成评估基准。其核心特点包括:(1)粗到细的任务设计,基于源图与参考图的存在与否,系统拆解为四类任务,进一步细化为31项细粒度任务,覆盖广泛生成需求;(2)多维度评估体系,涵盖美学质量、成像质量、提示遵循、源图一致性、参考图一致性与可控性六个维度,引入11项指标,其中创新提出VLLM-QA,利用大模型评估图像编辑成功率;(3)混合数据来源,结合真实场景与虚拟生成数据,提升数据多样性,缓解评估偏差。通过ICE-Bench对现有模型的全面分析揭示了当前能力与真实需求间的差距。为推动领域发展,我们将开源完整数据集、评估代码与模型资源。

原文摘要 · Abstract (English)

Image generation has witnessed significant advancements in the past few years. However, evaluating the performance of image generation models remains a formidable challenge. In this paper, we propose ICE-Bench, a unified and comprehensive benchmark designed to rigorously assess image generation models. Its comprehensiveness could be summarized in the following key features: (1) Coarse-to-Fine Tasks: We systematically deconstruct image generation into four task categories: No-ref/Ref Image Creating/Editing, based on the presence or absence of source images and reference images. And further decompose them into 31 fine-grained tasks covering a broad spectrum of image generation requirements, culminating in a comprehensive benchmark. (2) Multi-dimensional Metrics: The evaluation framework assesses image generation capabilities across 6 dimensions: aesthetic quality, imaging quality, prompt following, source consistency, reference consistency, and controllability. 11 metrics are introduced to support the multi-dimensional evaluation. Notably, we introduce VLLM-QA, an innovative metric designed to assess the success of image editing by leveraging large models. (3) Hybrid Data: The data comes from real scenes and virtual generation, which effectively improves data diversity and alleviates the bias problem in model evaluation. Through ICE-Bench, we conduct a thorough analysis of existing generation models, revealing both the challenging nature of our benchmark and the gap between current model capabilities and real-world generation requirements. To foster further advancements in the field, we will open-source ICE-Bench, including its dataset, evaluation code, and models, thereby providing a valuable resource for the research community.

图像生成评估基准多维度评价数据多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。