构建视频编辑理解基准,揭示大模型在此任务上的严重不足。
VEU-Bench: Towards Comprehensive Understanding of Video Editing
- 提出多维度视频编辑理解基准VEU-Bench,涵盖19项细粒度任务。
- 11个主流视频大模型在部分任务上表现低于随机猜测,准确率普遍偏低。
- 开发专用模型Oscars,性能超越开源模型28.3%,接近GPT-4o水平。
互联网上广泛传播的视频通常经过编辑。尽管视频大语言模型(Vid-LLMs)在通用视频理解任务上取得显著进展,其在视频编辑理解(VEU)任务上的能力仍不明确。为此,本文提出VEU-Bench(视频编辑理解基准),一个涵盖多种维度的综合性评估基准,从帧内特征如镜头大小到帧间属性如剪辑类型和转场方式。不同于以往仅关注编辑元素分类的基准,VEU-Bench包含三个阶段共19项细粒度任务:识别、推理与判断。为提升自动标注质量,我们构建了基于本体知识库的标注流水线。对11个前沿Vid-LLMs的实验表明,当前模型在VEU任务中面临严峻挑战,部分表现甚至低于随机选择。为此,我们训练了Oscars模型,该模型在精炼后的VEU-Bench数据集上微调,准确率超过现有开源模型28.3%,性能接近商业模型GPT-4o。此外,引入VEU数据可显著提升模型在通用视频理解基准上的表现,九项推理任务平均提升8.3%。
原文摘要 · Abstract (English)
Widely shared videos on the internet are often edited. Recently, although Video Large Language Models (Vid-LLMs) have made great progress in general video understanding tasks, their capabilities in video editing understanding (VEU) tasks remain unexplored. To address this gap, in this paper, we introduce VEU-Bench (Video Editing Understanding Benchmark), a comprehensive benchmark that categorizes video editing components across various dimensions, from intra-frame features like shot size to inter-shot attributes such as cut types and transitions. Unlike previous video editing understanding benchmarks that focus mainly on editing element classification, VEU-Bench encompasses 19 fine-grained tasks across three stages: recognition, reasoning, and judging. To enhance the annotation of VEU automatically, we built an annotation pipeline integrated with an ontology-based knowledge base. Through extensive experiments with 11 state-of-the-art Vid-LLMs, our findings reveal that current Vid-LLMs face significant challenges in VEU tasks, with some performing worse than random choice. To alleviate this issue, we develop Oscars, a VEU expert model fine-tuned on the curated VEU-Bench dataset. It outperforms existing open-source Vid-LLMs on VEU-Bench by over 28.3% in accuracy and achieves performance comparable to commercial models like GPT-4o. We also demonstrate that incorporating VEU data significantly enhances the performance of Vid-LLMs on general video understanding benchmarks, with an average improvement of 8.3% across nine reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。