无需训练,一步完成视频编辑,还解决时序闪烁问题。
ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

- 用共享噪声和运动对齐聚合帧间编辑场,实现时序一致
- 相比多步编辑,仅需2次前向传播,减少78%形变误差
- 适合追求高效、低延迟视频编辑的开发者与创作者
单步文本到图像模型可实现免训练、免反演的快速编辑,仅需1-2次网络前向传播(NFE)。ChordEdit通过采样时间上的低能量平滑稳定此类编辑。但独立应用于视频帧时,会引发时序闪烁和编辑强度漂移。本文提出ChordVideo,将低能量原则扩展至视频时间维度,采用共享噪声、运动对齐的因果聚合方式处理每帧的Chord场,并引入可选的时序平滑近端修正。理论推导出形变误差边界,分离运动偏差与随机闪烁,预测更大时间窗口下收益递减。在TGVE/DAVIS数据集上,使用两个单步骨干模型,ChordVideo使形变误差降低78%,闪烁减少49%,CLIP帧一致性提升9-10分,背景PSNR提高约1.5dB,且每帧保持2 NFE。相比七种多步编辑器,其在时序一致性和源内容保留上表现相当,但每片段模型步数减少10-60倍。
原文摘要 · Abstract (English)
One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes such edits through low-energy smoothing along sampling time. Applied independently to video frames, however, it produces temporal flicker and edit-strength drift. We introduce \textbf{ChordVideo}, which extends the same low-energy principle to video time through shared noise, motion-aligned causal aggregation of per-frame Chord fields, and an optional temporally smoothed proximal correction. We derive a warping-error bound that separates motion bias from stochastic flicker and predicts diminishing returns with larger temporal windows. On TGVE/DAVIS with two one-step backbones, ChordVideo reduces warping error by \textbf{78\%} and flicker by \textbf{49\%}, improves CLIP frame consistency by \textbf{9--10 points}, and increases background PSNR by about \textbf{1.5,dB}, while retaining \textbf{2 NFE/frame}. Compared with seven multi-step editors, it achieves competitive temporal consistency and source preservation using \textbf{10--60$\times$ fewer model steps per clip
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。