arXiv:2602.07095cs.CVcs.AI2026-02被引 2

构建首个面向世界知识的图像编辑数据集,提升模型理解隐含因果指令能力。

WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark

  • 设计基于真实因果逻辑的改写指令,引导模型理解隐含视觉变化原因
  • 引入因果验证奖励机制,在世界知识推理上接近GPT-4o表现
  • 适合研究开放世界图像编辑与常识推理的开发者和研究人员

近期图像编辑模型在执行明确指令(如属性修改、风格迁移、姿态合成)方面取得显著进展,但在处理隐含编辑指令时仍面临挑战——这类指令描述视觉变化的原因而非结果。现有模型依赖统一策略,缺乏对复杂世界知识与推理的需求支持。为此,我们提出 extbf{WorldEdit},一个专为世界驱动图像编辑设计的数据集,包含高质量编辑样本,其指令经过改写以符合真实因果逻辑。同时提供 extbf{WorldEdit-Test} 评估现有模型在因果编辑场景下的表现。采用两阶段训练框架微调如 Bagel 的模型,并引入因果验证奖励。实验表明,该方法显著缩小了与 GPT-4o 和 Nano-Banana 的差距,在指令遵循与知识合理性方面表现优异,尤其弥补了开源系统在常识推理上的短板。

原文摘要 · Abstract (English)

Recent advances in image editing models have demonstrated remarkable capabilities in executing explicit instructions, such as attribute manipulation, style transfer, and pose synthesis. However, these models often face challenges when dealing with implicit editing instructions, which describe the cause of a visual change without explicitly detailing the resulting outcome. These limitations arise because existing models rely on uniform editing strategies that are not equipped to handle the complex world knowledge and reasoning required for implicit instructions. To address this gap, we introduce \textbf{WorldEdit}, a dataset specifically designed to enable world-driven image editing. WorldEdit consists of high-quality editing samples, guided by paraphrased instructions that align with real-world causal logic. Furthermore, we provide \textbf{WorldEdit-Test} for evaluating the existing model's performance on causal editing scenarios. With WorldEdit, we use a two-stage training framework for fine-tuning models like Bagel, integrating with a causal verification reward. Our results show that the proposed dataset and methods significantly narrow the gap with GPT-4o and Nano-Banana, demonstrating competitive performance not only in instruction following but also in knowledge plausibility, where many open-source systems typically struggle.

图像编辑因果推理世界知识数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。