DAPE用两阶段微调实现高效视频编辑,兼顾时序一致性和视觉质量。
DAPE: Dual-Stage Parameter-Efficient Fine-Tuning for Consistent Video Editing with Diffusion Models
- 分两阶段:先优化时序一致性,再提升视觉质量
- 在多个数据集上显著提升视频连贯性与图文对齐效果
- 自建232视频大样本基准,解决旧数据集缺陷
基于扩散模型的视频生成是一项复杂的多模态任务,视频编辑成为该领域的重要方向。现有方法主要分为需训练和无需训练两类:前者计算成本高,后者性能不佳。为此,我们提出DAPE——一种高质量且低成本的两阶段参数高效微调框架。第一阶段设计了高效的归一化微调方法,增强生成视频的时序一致性;第二阶段引入视觉友好的适配器,提升视觉质量。此外,我们发现现有基准存在类别多样性不足、物体分布不均、帧数不一致等问题,因此构建了一个包含232个视频、6种编辑提示和丰富标注的大规模数据集基准,支持客观全面的评估。在BalanceCC、LOVEU-TGVE、RAVE等现有数据集及自建基准上的大量实验表明,DAPE显著提升了时序连贯性与文本-视频对齐能力,优于此前最先进方法。
原文摘要 · Abstract (English)
Video generation based on diffusion models presents a challenging multimodal task, with video editing emerging as a pivotal direction in this field. Recent video editing approaches primarily fall into two categories: training-required and training-free methods. While training-based methods incur high computational costs, training-free alternatives often yield suboptimal performance. To address these limitations, we propose DAPE, a high-quality yet cost-effective two-stage parameter-efficient fine-tuning (PEFT) framework for video editing. In the first stage, we design an efficient norm-tuning method to enhance temporal consistency in generated videos. The second stage introduces a vision-friendly adapter to improve visual quality. Additionally, we identify critical shortcomings in existing benchmarks, including limited category diversity, imbalanced object distribution, and inconsistent frame counts. To mitigate these issues, we curate a large dataset benchmark comprising 232 videos with rich annotations and 6 editing prompts, enabling objective and comprehensive evaluation of advanced methods. Extensive experiments on existing datasets (BalanceCC, LOVEU-TGVE, RAVE) and our proposed benchmark demonstrate that DAPE significantly improves temporal coherence and text-video alignment while outperforming previous state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。