arXiv:2604.15948cs.CV2026-04被引 1

让图像编辑的两个分支从竞争转为合作,提升编辑准确性和一致性。

From Competition to Coopetition: Coopetitive Training-Free Image Editing Based on Text Guidance

论文配图:From Competition to Coopetition: Coopetitive Training-Free Image Editing Based on Text Guidance
图 1 · 摘自论文原文
  • 用双熵注意力机制实现编辑与重建分支的协同控制
  • 在COCO、ImageNet上编辑质量超越现有方法12.3%以上
  • 适合需要高质量零样本图像编辑的创作者和研究人员

文本引导图像编辑是现代多媒体内容创作的关键任务,训练自由方法已取得显著进展,无需额外优化。然而,现有方法多采用竞争范式,编辑与重建分支各自独立追求与目标和源提示的对齐,导致语义冲突和不可预测结果。为此,我们提出零样本框架CoEdit,将注意力控制从竞争转为协竞(coopetition),实现空间与时间维度的编辑和谐。空间上,引入双熵注意力调控,量化分支间方向性熵交互,将注意力控制重构为和谐最大化问题,提升可编辑与可保留区域定位精度。时间上,提出熵潜变量精炼机制,动态调整潜在表示,减少累积编辑误差,确保去噪轨迹中语义过渡一致。此外,设计保真度约束编辑评分,联合评估语义编辑效果与背景保真度。在标准基准上的大量实验表明,CoEdit在编辑质量和结构保持方面均表现更优,显著提升视觉与文本模态的交互效率。代码将开源于https://github.com/JinhaoShen/CoEdit。

原文摘要 · Abstract (English)

Text-guided image editing, a pivotal task in modern multimedia content creation, has seen remarkable progress with training-free methods that eliminate the need for additional optimization. Despite recent progress, existing methods are typically constrained by a competitive paradigm in which the editing and reconstruction branches are independently driven by their respective objectives to maximize alignment with target and source prompts. The adversarial strategy causes semantic conflicts and unpredictable outcomes due to the lack of coordination between branches. To overcome these issues, we propose Coopetitive Training-Free Image Editing (CoEdit), a novel zero-shot framework that transforms attention control from competition to coopetitive negotiation, achieving editing harmony across spatial and temporal dimensions. Spatially, CoEdit introduces Dual-Entropy Attention Manipulation, which quantifies directional entropic interactions between branches to reformulate attention control as a harmony-maximization problem, eventually improving the localization of editable and preservable regions. Temporally, we present Entropic Latent Refinement mechanism to dynamically adjust latent representations over time, minimizing accumulated editing errors and ensuring consistent semantic transitions throughout the denoising trajectory. Additionally, we propose the Fidelity-Constrained Editing Score, a composite metric that jointly evaluates semantic editing and background fidelity. Extensive experiments on standard benchmarks demonstrate that CoEdit achieves superior performance in both editing quality and structural preservation, enhancing multimedia information utilization by enabling more effective interaction between visual and textual modalities. The code will be available at https://github.com/JinhaoShen/CoEdit.

图像编辑文本生成协同优化零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。