arXiv:2412.02114cs.CVcs.AI2024-12被引 2

无需标注数据,让视频生成模型直接实现通用编辑。

Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning

  • 通过自监督对齐,让生成模型同时掌握生成与编辑能力。
  • 仅需微调2.33%参数,即可在99个场景中实现跨类型编辑。
  • 适合想低成本提升视频模型编辑能力的研究者和开发者。

近期视频生成技术发展迅速,但视频编辑仍受限于三大问题:(a)依赖人工标注难以泛化;(b)生成与编辑任务人为割裂;(c)训练成本过高。本文提出轻量级自监督微调方法UES(Unlocking Universal Editing via Self-Supervision),将生成模型转化为统一的生成-编辑系统。通过原始视频-文本对构建双条件机制,联合提供视觉与语义信息,实现时空对应关系的结构化学习。优势包括:(i)无需监督即可适应多样编辑任务;(ii)适用于多数文本(+图像)到视频模型,统一生成与编辑流程;(iii)仅需微调2.33%参数(减少92.67%可调参数),显著提升效率。为系统评估,我们构建OmniBench-99基准,涵盖99个视频,覆盖人物、动物、环境、物体,包含4类编辑类型与8种场景。实验表明,即使无原生编辑能力的模型,经UES后也能实现强大且通用的编辑,同时保持或提升原有生成性能。

原文摘要 · Abstract (English)

Recent advances in video generation have outpaced progress in video editing, which remains constrained by several limiting factors, namely: (a) the task's dependency on supervision severely limits generality, (b) an unnecessary artificial separation between the generation and editing task, and (c) the high computational costs of training a video model. In this work, we propose UES (Unlocking Universal Editing via Self-Supervision), a lightweight self-supervised fine-tuning strategy that transforms generation models into unified generation-editing systems through self-supervised semantic alignment. Our approach establishes a dual-conditioning mechanism where original video-text pairs jointly provide visual and textual semantics, enabling structured learning of intrinsic spatiotemporal correspondences. Key advantages include: (i) Universality through supervision-free adaptation to diverse editing tasks, (ii) Unification of generation and editing applicable to most text(+image)-to-video model, and (iii) Efficiency via lightweight fine-tune that reduces tunable parameters by 92.67%. To enable systematic evaluation, we introduce OmniBench-99, a comprehensive benchmark spanning 99 videos across humans/animals, environments, and objects, comprising 4 editing types and 8 scenarios. Extensive experiments show UES enables models without inherent editing capability to perform powerful and universal editing while preserving or even enhancing their original generation performance.

视频编辑自监督学习生成模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。