arXiv:2505.01079cs.CVeess.IV2025-05CVPR被引 5

用分层记忆提升图像编辑的连续性,支持多次修改不破坏原有内容。

Improving Editability in Image Generation with Layer-wise Memory

  • 通过分层记忆存储历史编辑的潜在特征和提示嵌入。
  • 仅需粗略掩码即可完成多步编辑,保持场景一致性与内容质量。
  • 适合需要反复调整复杂图像的设计师或研究人员使用。

现实世界中的图像编辑通常需要多次连续操作才能达成理想效果。现有方法主要针对单对象修改设计,难以处理连续编辑:尤其在保留已有编辑结果的同时自然融入新元素方面表现不佳,严重限制了需同时修改多个对象且保持上下文关系的复杂场景。本文提出两种关键方案:支持粗略掩码输入以保留原有内容并自然融合新元素,以及实现多轮编辑的一致性。框架通过分层记忆机制存储先前编辑的潜在表示和提示嵌入,引入背景一致性引导利用记忆潜变量维持场景连贯性,并采用跨注意力中的多查询解耦确保新元素自然适配现有内容。为评估方法,我们构建了一个新基准数据集,包含语义对齐度量与交互式编辑场景。实验表明,在迭代编辑任务中表现优异,仅需粗略掩码,即可在多步操作中保持高质量输出。

原文摘要 · Abstract (English)

Most real-world image editing tasks require multiple sequential edits to achieve desired results. Current editing approaches, primarily designed for single-object modifications, struggle with sequential editing: especially with maintaining previous edits along with adapting new objects naturally into the existing content. These limitations significantly hinder complex editing scenarios where multiple objects need to be modified while preserving their contextual relationships. We address this fundamental challenge through two key proposals: enabling rough mask inputs that preserve existing content while naturally integrating new elements and supporting consistent editing across multiple modifications. Our framework achieves this through layer-wise memory, which stores latent representations and prompt embeddings from previous edits. We propose Background Consistency Guidance that leverages memorized latents to maintain scene coherence and Multi-Query Disentanglement in cross-attention that ensures natural adaptation to existing content. To evaluate our method, we present a new benchmark dataset incorporating semantic alignment metrics and interactive editing scenarios. Through comprehensive experiments, we demonstrate superior performance in iterative image editing tasks with minimal user effort, requiring only rough masks while maintaining high-quality results throughout multiple editing steps.

图像编辑分层记忆连续修改生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。