解决多轮图像编辑中身份漂移问题,保持长期一致性。
AnchorEdit: Maintaining Temporal Consistency in Multi-turn Image Editing via Causal Memory

- 用自回归扩散框架+因果记忆机制,逐轮稳定编辑
- 10轮以上编辑仍保持高主体保真度和指令遵循率
- 适合需要长时间迭代设计的AI绘画用户
多轮图像编辑对迭代设计至关重要,但现有模型在连续操作中常出现身份漂移和错误累积。尽管已有研究利用视频先验提升一致性,但其依赖双向注意力,与交互编辑的因果顺序本质不符。本文提出AnchorEdit,首个基于自回归扩散的高分辨率长时多轮编辑框架。通过三阶段训练流程:身份保持的单轮预训练、采用新型自回放策略的因果自回归微调以缓解暴露偏差、以及高效4步生成的一致性蒸馏。推理时引入记忆机制,锚定初始主体身份,确保长轨迹稳定外推。我们构建了一个新的高分辨率多轮编辑基准,用于测试长时稳定性。大量实验表明,AnchorEdit在10轮以上交互中仍保持优异主体保真度与指令遵循能力,达到当前最优效果。
原文摘要 · Abstract (English)
Multi-turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation over successive steps. While existing research leverages video priors for consistency, their reliance on bidirectional attention is fundamentally misaligned with the causal, sequential nature of interactive editing. In this paper, we propose AnchorEdit, the first autoregressive (AR) diffusion-based framework designed specifically for high-resolution, long-term multi-turn editing. AnchorEdit bridges the gap between video priors and causal inference through a three-stage training curriculum: identity-preserving sing-turn pretraining, causal AR forcing fine-tuning with a novel self-rollout strategy to mitigate exposure bias, and consistency distillation for efficient 4-step generation. During inference, we introduce a memory mechanism to anchor the initial subject identity and ensure stable extrapolation across extended editing trajectories. To evaluate performance, we provide a new high-resolution multi-turn editing benchmark designed to stress-test long-horizon stability. Extensive experiments demonstrate that AnchorEdit achieves state-of-the-art results, maintaining exceptional subject fidelity and instruction following even over 10+ interaction rounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。