构建30万条高清视频对,实现无需实时引导的精准视频编辑。
FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing
- 基于30万条720p长视频对,训练出能自主推断运动与外观的模型。
- 在编辑基准上提升0.2的PickScore和0.3的VLM评分,显著超越现有方法。
- 适合需要高稳定性、无交互式编辑的视频生成与修复场景。
首帧传播(FFP)为可控视频编辑提供了新范式,但现有方法依赖繁琐的运行时引导。我们发现其根源在于训练数据不足:当前数据集通常过短、分辨率低,且任务多样性缺失,难以建立鲁棒的时间先验。为此,我们提出首个大规模数据集FFP-300K,包含30万对720p分辨率、每段81帧的高质量视频对,通过双轨化流程覆盖局部与全局编辑。基于此,我们设计了真正无引导的FFP框架,解决保持首帧外观与保留源视频运动之间的关键矛盾。架构上引入自适应时空位置编码(AST-RoPE),动态重映射位置信息以解耦外观与运动表征;目标层面采用自蒸馏策略,以身份传播任务作为强正则化项,保障长期时间稳定性并防止语义漂移。在EditVerseBench基准上的全面实验表明,本方法相比现有学术与商业模型,在性能上分别提升约0.2的PickScore和0.3的VLM分数。
原文摘要 · Abstract (English)
First-Frame Propagation (FFP) offers a promising paradigm for controllable video editing, but existing methods are hampered by a reliance on cumbersome run-time guidance. We identify the root cause of this limitation as the inadequacy of current training datasets, which are often too short, low-resolution, and lack the task diversity required to teach robust temporal priors. To address this foundational data gap, we first introduce FFP-300K, a new large-scale dataset comprising 300K high-fidelity video pairs at 720p resolution and 81 frames in length, constructed via a principled two-track pipeline for diverse local and global edits. Building on this dataset, we propose a novel framework designed for true guidance-free FFP that resolves the critical tension between maintaining first-frame appearance and preserving source video motion. Architecturally, we introduce Adaptive Spatio-Temporal RoPE (AST-RoPE), which dynamically remaps positional encodings to disentangle appearance and motion references. At the objective level, we employ a self-distillation strategy where an identity propagation task acts as a powerful regularizer, ensuring long-term temporal stability and preventing semantic drift. Comprehensive experiments on the EditVerseBench benchmark demonstrate that our method significantly outperforming existing academic and commercial models by receiving about 0.2 PickScore and 0.3 VLM score improvement against these competitors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。