提出视频差异描述新任务,让模型精准对比两段视频的异同。
ViDiC: Video Difference Captioning
- 设计视频对比任务与双检查表评估框架,区分相似与差异判断。
- 构建含1000对视频的ViDiC-1K数据集,涵盖7类差异维度。
- 揭示主流多模态模型在视频对比理解上存在显著能力差距。
理解动态场景间的视觉差异需要对构图、空间和时间变化进行比较感知,这一能力在现有视觉语言系统中仍待探索。尽管图像差异描述(IDC)已有研究,使模型能描述静态图像间的语义变化,但现有方法无法捕捉运动连续性、事件演变或编辑一致性。本文提出视频差异描述(ViDiC)任务及对应的ViDiC-1K数据集,用于评估多模态大模型对视频对间细粒度异同的描述能力。该数据集包含1,000组精心筛选的视频对,标注超过4,000个对比检查项,覆盖主体、风格、背景、摄法、运动、位置和播放技术七类差异。为确保评估可靠性,提出基于大模型作为评判者(LLM-as-a-Judge)的双检查表框架,分别衡量相似性与差异性的准确率。在19个代表性多模态模型上的实验显示,其对比描述与差异感知能力存在显著差距。我们期望ViDiC-1K能成为推动视频理解、编辑意识与比较推理发展的关键基准。
原文摘要 · Abstract (English)
Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored in existing vision-language systems. While prior work on Image Difference Captioning (IDC) has enabled models to describe semantic changes between static images, these approaches fail to capture motion continuity, event evolution, or editing consistency over time. We introduce the ViDiC (Video Difference Captioning) task and its corresponding ViDiC-1K dataset, designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to provide fine-grained descriptions of similarities and differences between video pairs. ViDiC-1K comprises 1,000 curated video pairs annotated with over 4,000 comparative checklist items, covering seven categories: subject, style, background, cinematography, motion, location, and playback techniques. To ensure reliable evaluation, we propose a dual-checklist framework that measures the accuracy of similarity and difference separately, based on the LLM-as-a-Judge protocol. Experiments on nineteen representative multimodal models reveal a significant performance gap in their comparative description and difference perception abilities. We hope ViDiC-1K can be a challenging benchmark that lays a solid foundation for advancing video understanding, edit awareness, and comparative reasoning in multimodal intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。