arXiv:2511.01295cs.CV2025-11被引 20

构建首个面向推理的图像编辑统一评测基准,覆盖真实与游戏世界场景。

UniREditBench: A Unified Reasoning-based Image Editing Benchmark

  • 设计多场景、多维度的2700个样本,涵盖8大类18子类编辑任务。
  • 引入图文双参考评估机制,提升复杂推理任务评价可靠性。
  • 开源10万条带思维链标注的合成数据,适配研究者与工业界模型评测。

近年来多模态生成模型在图像编辑领域取得显著进展,但现有模型在需要隐式推理的复杂编辑任务中仍表现不足,亟需一个系统性评测基准。当前基准主要聚焦于真实场景下单对象属性变换,存在两大缺陷:一是忽视多对象交互及基于规则的游戏世界场景;二是仅依赖文本参考评估,易在复杂推理中产生误判。为此,本文提出UniREditBench,一个面向推理的统一图像编辑评测基准,包含2,700个精心构建的样本,覆盖真实世界与游戏世界场景,涵盖8个主维度和18个子维度。为提升评估可靠性,引入图文双参考评估机制,提供文本与真实图像双重参考。此外,设计自动化多场景数据合成流水线,构建大规模合成数据集UniREdit-Data-100K,包含高质量链式思维(CoT)标注。基于此数据微调Bagel模型,开发UniREdit-Bagel,在域内与域外设置均表现显著提升。通过全面评测开源与闭源模型,揭示其在各类任务中的优劣势。

原文摘要 · Abstract (English)

Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still struggle with handling diverse and complex image editing tasks that require implicit reasoning, underscoring the need for a comprehensive benchmark to systematically assess their performance across various reasoning scenarios. Existing benchmarks primarily focus on single-object attribute transformation in realistic scenarios, which, while effective, encounter two key challenges: (1) they largely overlook multi-object interactions as well as game-world scenarios that involve human-defined rules, which are common in real-life applications; (2) they only rely on textual references to evaluate the generated images, potentially leading to systematic misjudgments, especially in complex reasoning scenarios. To this end, this work proposes UniREditBench, a unified benchmark for reasoning-based image editing evaluation. It comprises 2,700 meticulously curated samples, covering both real- and game-world scenarios across 8 primary dimensions and 18 sub-dimensions. To improve evaluation reliability, we introduce multimodal dual-reference evaluation, providing both textual and ground-truth image references for each sample assessment. Furthermore, we design an automated multi-scenario data synthesis pipeline and construct UniREdit-Data-100K, a large-scale synthetic dataset with high-quality chain-of-thought (CoT) reasoning annotations. We fine-tune Bagel on this dataset and develop UniREdit-Bagel, demonstrating substantial improvements in both in-domain and out-of-distribution settings. Through thorough benchmarking of both open-source and closed-source image editing models, we reveal their strengths and weaknesses across various aspects.

图像编辑推理评测多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。