arXiv:2603.05518cs.HCcs.CV2026-03被引 1

通过认知分步推理实现更精准的自然语言图像编辑

CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning

  • 将编辑任务拆解为‘改什么’和‘怎么改’,分两阶段进行认知推理
  • 在通用与隐私合规评测中均达开源模型最优,视觉一致性显著提升
  • 无需训练、全开源组件构建,适合追求可解释性与透明性的研究者

大型多模态模型的发展使基于指令的图像编辑成为可能,用户可通过自然语言描述修改视觉内容。然而,现有方法在处理模糊或复杂的指令时,常面临高层次语义推理与视觉一致性不足的问题。为此,我们提出 CoEditor++,一种基于认知结构、无需训练的框架,通过两个认知阶段与反射式自选择机制,将编辑分解为‘改什么’和‘怎么改’,实现鲁棒、精细且可解释的编辑。该框架完全由开源组件构成,无需额外训练或微调,确保透明性与跨领域适用性。我们在 SmartEdit(通用编辑基准)和 AltBear(隐私与合规导向基准)上评估了 CoEditor++,结果表明其在通用编辑与负责任编辑任务中均优于需训练的开源模型,且在视觉一致性方面显著领先。与 Nano Banana Pro、GPT-4o 等闭源模型相比,CoEditor++ 在指令遵循能力相当的前提下,视觉一致性明显更优。消融实验验证了其有效性源于结构化认知设计,而非特定组件。研究提示了以认知为中心的指令式图像编辑的潜力。

原文摘要 · Abstract (English)

Recent advances in large multimodal models (LMMs) have enabled instruction-based image editing, allowing users to modify visual content via natural language descriptions. However, existing approaches often struggle with high-level semantic reasoning and visual consistency, particularly under ambiguous or complex instructions. To address these challenges, we propose CoEditor++, a cognitively structured, training-free framework that decomposes editing into "what to edit" and "how to edit" through two cognitive stages with a reflective self-selection mechanism, enabling robust, fine-grained, and interpretable editing. Built entirely from open-sourced components, CoEditor++ requires no additional training or fine-tuning, ensuring transparency and cross-domain applicability. We evaluate CoEditor++ on SmartEdit, a widely used benchmark for general editing, and AltBear, a privacy and compliance-oriented benchmark. Experimental results show that CoEditor++ achieves state-of-the-art performance in both general editing and responsible editing tasks compared with open-sourced models that require training on specialized editing datasets maintaining significantly higher visual consistency. When compared with closed-source models such as Nano Banana Pro or GPT-4o, CoEditor++ preserves comparable instruction following while still substantially outperforming them in visual consistency. Extensive ablation studies confirm that the effectiveness of CoEditor++ benefits from its structured cognitive design rather than any specific model component. Our findings suggest the potential toward cognitive-centric instruction-based image editing.

图像编辑多模态认知推理开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。