arXiv:2503.14935cs.CVcs.AI2025-03NeurIPS被引 25

构建视频细粒度运动理解新基准,揭示大模型短板

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

  • 设计1776个带精细标注的视频数据集,覆盖闭合与开放任务
  • 21个顶尖多模态模型在细粒度运动理解上表现普遍不足
  • 配套训练数据集可提升模型对动态细节的捕捉能力

多模态大语言模型在视频内容理解方面展现出强大能力,但在细粒度运动理解上仍存在明显不足。为全面评估现有模型的运动理解能力,我们提出了FAVOR-Bench,包含1,776个视频和结构化的人工标注运动信息。该基准涵盖闭合式与开放式任务:闭合式评估设计了8,184个多项选择题,覆盖六个子任务;开放式评估提出一种新型低成本、无需LLM的评估方法,以及基于GPT的辅助评估方法,提升可解释性与可复现性。对21个前沿多模态大模型的综合实验表明,它们在描述视频中详细时序动态方面存在显著局限。为进一步缓解此问题,我们构建了包含17,152个视频的FAVOR-Train数据集,用于微调。以Qwen2.5-VL在FAVOR-Train上微调后,在TVBench、MotionBench及FAVOR-Bench的运动相关任务上均取得一致提升。结果表明,FAVOR-Bench与FAVOR-Train为社区开发更强视频理解模型提供了重要工具。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown remarkable capabilities in video content understanding but still struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding ability of existing MLLMs, we introduce FAVOR-Bench, comprising 1,776 videos with structured manual annotations of various motions. Our benchmark includes both close-ended and open-ended tasks. For close-ended evaluation, we carefully design 8,184 multiple-choice question-answer pairs spanning six distinct sub-tasks. For open-ended evaluation, we develop both a novel cost-efficient LLM-free and a GPT-assisted caption assessment method, where the former can enhance benchmarking interpretability and reproducibility. Comprehensive experiments with 21 state-of-the-art MLLMs reveal significant limitations in their ability to comprehend and describe detailed temporal dynamics in video motions. To alleviate this limitation, we further build FAVOR-Train, a dataset consisting of 17,152 videos with fine-grained motion annotations. The results of finetuning Qwen2.5-VL on FAVOR-Train yield consistent improvements on motion-related tasks of TVBench, MotionBench and our FAVOR-Bench. Comprehensive assessment results demonstrate that the proposed FAVOR-Bench and FAVOR-Train provide valuable tools to the community for developing more powerful video understanding models. Project page: \href{https://favor-bench.github.io/}{https://favor-bench.github.io/}.

视频理解运动分析多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。