arXiv:2603.16944cs.CVcs.AI2026-03被引 1

新基准测试发现图像编辑模型在复杂任务中表现不稳定。

Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models

  • 设计双轨评测体系,区分单轮一致性和多轮协同性
  • 8个主流模型在高语义任务上性能普遍下降
  • 专为专业场景设计,适合图像编辑系统开发者

尽管基于指令的图像编辑(IIE)取得显著进展,现有基准多采用混合评估方式,掩盖了专业应用中关键的性能不一致性问题:模型在不同语义尺度任务上的表现波动。为此,我们提出Omni IIE Bench,一个高质量、人工标注的基准,专门用于诊断IIE模型在实际应用中的编辑一致性。该基准采用创新的双轨诊断设计:(1) 单轮一致性,包含属性修改与实体替换的共享上下文任务对;(2) 多轮协调,涵盖跨越语义尺度的连续对话任务。基准通过严格的多阶段人工筛选流程构建,标准由计算机视觉研究生执行,行业相关性由专业设计师评审。我们对8个主流IIE模型进行了全面评估,首次量化揭示普遍存在的性能差距:几乎所有模型在从低语义尺度过渡到高语义尺度任务时均出现显著性能下降。Omni IIE Bench为下一代更可靠、稳定的IIE模型开发提供了关键诊断工具与洞察。

原文摘要 · Abstract (English)

While Instruction-based Image Editing (IIE) has achieved significant progress, existing benchmarks pursue task breadth via mixed evaluations. This paradigm obscures a critical failure mode crucial in professional applications: the inconsistent performance of models across tasks of varying semantic scales. To address this gap, we introduce Omni IIE Bench, a high-quality, human-annotated benchmark specifically designed to diagnose the editing consistency of IIE models in practical application scenarios. Omni IIE Bench features an innovative dual-track diagnostic design: (1) Single-turn Consistency, comprising shared-context task pairs of attribute modification and entity replacement; and (2) Multi-turn Coordination, involving continuous dialogue tasks that traverse semantic scales. The benchmark is constructed via an exceptionally rigorous multi-stage human filtering process, incorporating a quality standard enforced by computer vision graduate students and an industry relevance review conducted by professional designers. We perform a comprehensive evaluation of 8 mainstream IIE models using Omni IIE Bench. Our analysis quantifies, for the first time, a prevalent performance gap: nearly all models exhibit a significant performance degradation when transitioning from low-semantic-scale to high-semantic-scale tasks. Omni IIE Bench provides critical diagnostic tools and insights for the development of next-generation, more reliable, and stable IIE models.

图像编辑基准测试一致性AI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。