为文本驱动视频编辑设计了新评估体系,包含数据集与专用评分模型。
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
- 构建大规模数据集TDVE-DB,涵盖12种模型、8类编辑任务。
- 提出TDVE-Assessor模型,在3个维度上超越现有方法。
- 适合视频生成、评估算法研究者使用,推动领域标准化。
文本驱动视频编辑快速发展,但其严格评估仍面临挑战,因缺乏能捕捉编辑质量细微差别的专用视频质量评估(VQA)模型。为填补这一空白,我们提出TDVE-DB,一个大规模基准数据集,包含3,857个由12种不同模型生成的编辑视频,覆盖8类编辑任务,并附有173,565条人工主观评分,涵盖编辑质量、编辑对齐度和结构一致性三个关键维度。基于该数据集,我们全面评估了12个先进编辑模型的性能,揭示当前技术的优劣,并对现有VQA方法在文本驱动视频编辑评估中的表现进行基准测试。在此基础上,我们提出TDVE-Assessor,一种专为文本驱动视频编辑设计的新型VQA模型,将空间与时间视频特征融合至大语言模型(LLM),实现丰富上下文理解,提供全面质量评估。大量实验表明,TDVE-Assessor在所有三个评估维度上均显著优于现有VQA模型,在TDVE-DB上达到新的最先进水平。TDVE-DB与TDVE-Assessor将在论文发表后公开。
原文摘要 · Abstract (English)
Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. To address this critical gap, we introduce TDVE-DB, a large-scale benchmark dataset for text-driven video editing. TDVE-DB consists of 3,857 edited videos generated from 12 diverse models across 8 editing categories, and is annotated with 173,565 human subjective ratings along three crucial dimensions, i.e., edited video quality, editing alignment, and structural consistency. Based on TDVE-DB, we first conduct a comprehensive evaluation for the 12 state-of-the-art editing models revealing the strengths and weaknesses of current video techniques, and then benchmark existing VQA methods in the context of text-driven video editing evaluation. Building on these insights, we propose TDVE-Assessor, a novel VQA model specifically designed for text-driven video editing assessment. TDVE-Assessor integrates both spatial and temporal video features into a large language model (LLM) for rich contextual understanding to provide comprehensive quality assessment. Extensive experiments demonstrate that TDVE-Assessor substantially outperforms existing VQA models on TDVE-DB across all three evaluation dimensions, setting a new state-of-the-art. Both TDVE-DB and TDVE-Assessor will be released upon the publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。