arXiv:2512.01755cs.CV2025-12被引 4

解决多轮图像编辑中细节丢失问题,保持图像清晰度。

FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image Editing

  • 通过注入参考速度场的高频特征,保留图像细粒度细节。
  • 在10轮以上连续编辑中,身份保持和指令遵循能力显著优于现有方法。
  • 无需训练,适合需要稳定多轮修改的视觉创作场景。

基于自然语言的图像编辑已成为直观视觉操作的强大范式。尽管近期模型在单次编辑上表现优异,但在多轮编辑中会严重退化。系统分析表明,高频信息逐渐丢失是主要原因。我们提出 FreqEdit,一种无需训练的框架,可在10次以上连续迭代中保持稳定编辑。其包含三个协同组件:(1) 从参考速度场注入高频特征以保留细粒度细节;(2) 自适应注入策略,空间调节注入强度以实现区域精准控制;(3) 路径补偿机制,周期性重校正编辑轨迹,防止过度约束。大量实验表明,与七种先进基线相比,FreqEdit在身份保持和指令遵循方面均表现更优。

原文摘要 · Abstract (English)

Instruction-based image editing through natural language has emerged as a powerful paradigm for intuitive visual manipulation. While recent models achieve impressive results on single edits, they suffer from severe quality degradation under multi-turn editing. Through systematic analysis, we identify progressive loss of high-frequency information as the primary cause of this quality degradation. We present FreqEdit, a training-free framework that enables stable editing across 10+ consecutive iterations. Our approach comprises three synergistic components: (1) high-frequency feature injection from reference velocity fields to preserve fine-grained details, (2) an adaptive injection strategy that spatially modulates injection strength for precise region-specific control, and (3) a path compensation mechanism that periodically recalibrates the editing trajectory to prevent over-constraint. Extensive experiments demonstrate that FreqEdit achieves superior performance in both identity preservation and instruction following compared to seven state-of-the-art baselines.

图像编辑多轮操作高频保留生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。