arXiv:2510.11647cs.CV2025-10中稿 · ICLR被引 27

构建首个专用于指令引导视频编辑的综合性评测基准

IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment

  • 设计包含600个高质量视频的多样数据集,覆盖7类语义维度
  • 涵盖8类编辑任务35个子类,通过大模型与专家优化提示
  • 三维度评估体系融合质量、指令遵循与保真度,支持人机对齐

指令引导视频编辑正快速发展,为内容创作带来新机遇,但也面临系统性评估挑战。现有评测基准在源数据多样性、任务覆盖范围和评估指标完整性方面存在不足。为此,我们提出IVEBench——一个专为指令引导视频编辑设计的现代评测套件。它包含600个高质量源视频,覆盖7个语义维度,视频长度介于32至1,024帧之间;涵盖8类编辑任务、35个子类别,提示由大语言模型生成并经专家评审优化。关键在于,IVEBench建立了一个包含视频质量、指令遵循度与视频保真度的三维评估协议,融合传统指标与多模态大模型评估。大量实验表明,该套件能有效评测先进方法,提供全面且符合人类判断的评估结果。

原文摘要 · Abstract (English)

Instruction-guided video editing has emerged as a rapidly advancing research direction, offering new opportunities for intuitive content transformation while also posing significant challenges for systematic evaluation. Existing video editing benchmarks fail to support the evaluation of instruction-guided video editing adequately and further suffer from limited source diversity, narrow task coverage and incomplete evaluation metrics. To address the above limitations, we introduce IVEBench, a modern benchmark suite specifically designed for instruction-guided video editing assessment. IVEBench comprises a diverse database of 600 high-quality source videos, spanning seven semantic dimensions, and covering video lengths ranging from 32 to 1,024 frames. It further includes 8 categories of editing tasks with 35 subcategories, whose prompts are generated and refined through large language models and expert review. Crucially, IVEBench establishes a three-dimensional evaluation protocol encompassing video quality, instruction compliance and video fidelity, integrating both traditional metrics and multimodal large language model-based assessments. Extensive experiments demonstrate the effectiveness of IVEBench in benchmarking state-of-the-art instruction-guided video editing methods, showing its ability to provide comprehensive and human-aligned evaluation outcomes.

视频编辑评测基准指令引导多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。