系统梳理视频大模型评估基准,指出现有方法短板与改进方向。
VideoLLM Benchmarks and Evaluation: A Survey
- 分析闭集、开集及时空理解专用评估范式
- 揭示当前顶尖视频大模型在多基准上的性能趋势
- 适合关注视频理解评估体系的研者参考
大型语言模型(LLMs)的快速发展推动了视频理解技术的重大进步。本综述全面分析了专为视频大语言模型(VideoLLMs)设计或使用的基准测试与评估方法。文章探讨了当前视频理解基准的特征、评估协议及其局限性,分析了闭集、开集以及针对时间与时空理解任务的专项评估方法。研究揭示了前沿VideoLLMs在各类基准上的表现趋势,并识别出现有评估框架中的关键挑战。此外,提出未来研究方向,包括构建更多样、多模态和可解释性的评估基准,优化评估指标与流程。本综述旨在帮助研究人员系统掌握VideoLLMs的有效评估方法,明确推动视频理解领域发展的潜在路径。
原文摘要 · Abstract (English)
The rapid development of Large Language Models (LLMs) has catalyzed significant advancements in video understanding technologies. This survey provides a comprehensive analysis of benchmarks and evaluation methodologies specifically designed or used for Video Large Language Models (VideoLLMs). We examine the current landscape of video understanding benchmarks, discussing their characteristics, evaluation protocols, and limitations. The paper analyzes various evaluation methodologies, including closed-set, open-set, and specialized evaluations for temporal and spatiotemporal understanding tasks. We highlight the performance trends of state-of-the-art VideoLLMs across these benchmarks and identify key challenges in current evaluation frameworks. Additionally, we propose future research directions to enhance benchmark design, evaluation metrics, and protocols, including the need for more diverse, multimodal, and interpretability-focused benchmarks. This survey aims to equip researchers with a structured understanding of how to effectively evaluate VideoLLMs and identify promising avenues for advancing the field of video understanding with large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。