arXiv:2602.02123cs.CV2026-02

无需训练,高效编辑分钟级视频,解决长视频抖动与结构漂移问题

MLV-Edit: Towards Consistent and Highly Efficient Editing for Minute-Level Videos

论文配图:MLV-Edit: Towards Consistent and Highly Efficient Editing for Minute-Level Videos
图 1 · 摘自论文原文
  • 分段处理+运动场对齐,消除片段边界抖动
  • 全局参考锚定局部特征,抑制长期结构漂移
  • 适合需要稳定长视频编辑的创作者与开发者

我们提出 MLV-Edit,一种无需训练的基于流的框架,针对分钟级视频编辑的独特挑战。现有方法在短视频操作中表现优异,但扩展到长时间视频时面临计算开销过大及跨数千帧保持时间一致性困难的问题。MLV-Edit 采用分而治之策略进行分段编辑,包含两个核心模块:Velocity Blend 通过对齐相邻片段的光流场,修正段间运动不一致,消除常见于分段处理中的闪烁和边界伪影;Attention Sink 将局部片段特征锚定于全局参考帧,有效抑制累积的结构漂移。大量定量与定性实验表明,MLV-Edit 在时间稳定性与语义保真度方面均持续优于当前最先进方法。

原文摘要 · Abstract (English)

We propose MLV-Edit, a training-free, flow-based framework that address the unique challenges of minute-level video editing. While existing techniques excel in short-form video manipulation, scaling them to long-duration videos remains challenging due to prohibitive computational overhead and the difficulty of maintaining global temporal consistency across thousands of frames. To address this, MLV-Edit employs a divide-and-conquer strategy for segment-wise editing, facilitated by two core modules: Velocity Blend rectifies motion inconsistencies at segment boundaries by aligning the flow fields of adjacent chunks, eliminating flickering and boundary artifacts commonly observed in fragmented video processing; and Attention Sink anchors local segment features to global reference frames, effectively suppressing cumulative structural drift. Extensive quantitative and qualitative experiments demonstrate that MLV-Edit consistently outperforms state-of-the-art methods in terms of temporal stability and semantic fidelity.

视频编辑长视频流模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。