arXiv:2606.14042cs.CV2026-06

揭示单步图像编辑的内在机制,提出动态时间步自适应编辑新思路

Rethinking One-Step Image Editing through ChordEdit: Reproduction, Simplification, and New Insights

论文配图:Rethinking One-Step Image Editing through ChordEdit: Reproduction, Simplification, and New Insights
图 1 · 摘自论文原文
  • 将编辑分解为低频粗调与高频精修两阶段
  • 发现时间偏移δ是关键调控因子,影响编辑效果
  • 适合需要快速、精准文本引导编辑的研究者

单步图像编辑对实现快速、实用且易部署的文本引导编辑至关重要,但其内在机制仍不明确。我们通过复现、消融实验和简化分析重新审视ChordEdit。结果表明:a) 音程窗口δ主要作为从t到t−δ的有效时间步偏移;b) 音程传输作用于高噪声图像,主要执行低频语义编辑;c) 近邻对齐作用于低噪声图像,补充添加高频目标细节。在此视角下,ChordEdit自然分解为粗粒度低频传输阶段与细粒度高频对齐阶段。这些发现提示了基于提示的动态时间步选择路径,可实现自适应图像编辑。所有代码与结果见https://github.com/Harvard-AI-and-Robotics-Lab/ChordEdit-Reproduction。

原文摘要 · Abstract (English)

One-step image editing is important for making text-guided editing fast, practical, and easy to deploy, but its underlying mechanism is still not fully understood. We revisit ChordEdit through reproduction, ablation, and simplification. Our analysis shows that a) the chord window $δ$ largely acts as an effective timestep shift from $t$ to $t - δ$; b) chord transport acts on high-noise images and mainly performs low-frequency semantic editing; and c) proximal alignment acts on low-noise images and complements it by adding high-frequency target details. In this view, ChordEdit naturally decomposes editing into a coarse low-frequency transport stage and a fine high-frequency alignment stage. These findings suggest a path toward prompt-conditioned dynamic timestep selection for adaptive image editing. All code and results can be found at \href{https://github.com/Harvard-AI-and-Robotics-Lab/ChordEdit-Reproduction}{link}.

图像编辑扩散模型时间步控制文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。