arXiv:2512.04660cs.CV2025-12被引 7

构建首个覆盖10类任务的图像编辑评测基准,自动评估30个细粒度维度。

I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models

  • 设计涵盖单图与多图编辑的10类任务,扩展评测边界
  • 集成自动化工具与大模型,实现30个细分维度的精准评估
  • 通过人类偏好验证,确保评测结果真实可信

图像编辑模型发展迅速,但全面评估仍面临挑战。现有评测基准普遍存在任务范围有限、评估维度不足、依赖人工标注等问题,严重制约其可扩展性与实用性。为此,我们提出I2I-Bench,一个全面的图像到图像编辑模型评测基准,包含:(i) 多样化任务,涵盖单图与多图编辑的10类任务;(ii) 全面评估维度,包含30个解耦且细粒度的评估指标,采用自动化混合评估方法,融合专用工具与大型多模态模型(LMMs);(iii) 严格的对齐验证,证明基准评估结果与人类偏好高度一致。基于I2I-Bench,我们对多个主流图像编辑模型进行了评测,揭示了不同模型在各维度上的差距与权衡。所有组件将开源,以推动后续研究。

原文摘要 · Abstract (English)

Image editing models are advancing rapidly, yet comprehensive evaluation remains a significant challenge. Existing image editing benchmarks generally suffer from limited task scopes, insufficient evaluation dimensions, and heavy reliance on manual annotations, which significantly constrain their scalability and practical applicability. To address this, we propose \textbf{I2I-Bench}, a comprehensive benchmark for image-to-image editing models, which features (i) diverse tasks, encompassing 10 task categories across both single-image and multi-image editing tasks, (ii) comprehensive evaluation dimensions, including 30 decoupled and fine-grained evaluation dimensions with automated hybrid evaluation methods that integrate specialized tools and large multimodal models (LMMs), and (iii) rigorous alignment validation, justifying the consistency between our benchmark evaluations and human preferences. Using I2I-Bench, we benchmark numerous mainstream image editing models, investigating the gaps and trade-offs between editing models across various dimensions. We will open-source all components of I2I-Bench to facilitate future research.

图像编辑评测基准自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。