arXiv:2503.13684cs.CV2025-03被引 15

提出细粒度视频编辑基准FiVE,评测生成模型的精准修改能力。

FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models

  • 构建包含74实拍+26生成视频的细粒度编辑数据集
  • 基于流模型实现无需训练的高效视频编辑,性能优于扩散模型
  • 引入VLM评估新指标,更准确衡量对象级编辑成功率

近期涌现大量文本到视频编辑方法,但缺乏标准化评估基准导致评价不一致且难分析超参数敏感性。细粒度视频编辑对实现物体级精准修改、保持上下文与时间一致性至关重要。为此,我们提出FiVE——一个用于评估新兴扩散与修正流模型的细粒度视频编辑基准。该基准包含74个真实世界视频和26个生成视频,涵盖6类细粒度编辑任务、420组物体级编辑提示对及其对应掩码。我们通过引入FlowEdit,将最新修正流(RF)T2V生成模型Pyramid-Flow和Wan2.1改造为无需训练和反演的视频编辑模型Pyramid-Edit和Wan-Edit。在FiVE基准上,使用15项指标(涵盖背景保留、文本-视频相似性、时间一致性、视频质量与运行时)评估五种基于扩散、两种基于修正流的编辑方法。为进一步提升物体级评估能力,我们提出FiVE-Acc,一种利用视觉语言模型(VLMs)的新指标,用于判断细粒度编辑的成功率。实验结果表明,基于修正流的编辑方法显著优于基于扩散的方法,其中Wan-Edit表现最佳,且对超参数最不敏感。更多视频演示见匿名网站:https://sites.google.com/view/five-benchmark

原文摘要 · Abstract (English)

Numerous text-to-video (T2V) editing methods have emerged recently, but the lack of a standardized benchmark for fair evaluation has led to inconsistent claims and an inability to assess model sensitivity to hyperparameters. Fine-grained video editing is crucial for enabling precise, object-level modifications while maintaining context and temporal consistency. To address this, we introduce FiVE, a Fine-grained Video Editing Benchmark for evaluating emerging diffusion and rectified flow models. Our benchmark includes 74 real-world videos and 26 generated videos, featuring 6 fine-grained editing types, 420 object-level editing prompt pairs, and their corresponding masks. Additionally, we adapt the latest rectified flow (RF) T2V generation models, Pyramid-Flow and Wan2.1, by introducing FlowEdit, resulting in training-free and inversion-free video editing models Pyramid-Edit and Wan-Edit. We evaluate five diffusion-based and two RF-based editing methods on our FiVE benchmark using 15 metrics, covering background preservation, text-video similarity, temporal consistency, video quality, and runtime. To further enhance object-level evaluation, we introduce FiVE-Acc, a novel metric leveraging Vision-Language Models (VLMs) to assess the success of fine-grained video editing. Experimental results demonstrate that RF-based editing significantly outperforms diffusion-based methods, with Wan-Edit achieving the best overall performance and exhibiting the least sensitivity to hyperparameters. More video demo available on the anonymous website: https://sites.google.com/view/five-benchmark

视频编辑扩散模型修正流评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。