arXiv:2411.13281cs.CVcs.AI2024-11CVPR被引 19

用自动化用户模拟评估大模型视频理解能力,更贴近真实需求。

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

  • 通过用户模拟生成动态开放问题,替代传统选择题评测
  • 采用改进的ELO系统实现多模型持续公平对比,结果与人工判断高度一致
  • 适合关注视频分析模型真实表现、追求高效评测的研究者

具备先进视频分析能力的大规模多模态模型(LMMs)近年来备受关注。然而,现有评估大多依赖如VideoMME和LongVideoBench等基准中的传统多选题,难以捕捉真实用户复杂需求。为解决这一问题——同时克服人工标注成本高、速度慢的瓶颈——我们提出VideoAutoArena,一个受LMSYS Chatbot Arena启发的竞技场式评测框架,可自动评估LMM在视频理解方面的能力。该框架利用用户模拟生成开放式、自适应的问题,全面考察模型表现,并引入改进的ELO评分系统,实现多个LMM间的公平、持续比较。为验证其有效性,我们构建了基于人工标注的“黄金标准”,证明该评测体系与人类判断高度一致且可扩展。此外,我们提出故障驱动演化策略,逐步提升问题难度,推动模型应对更复杂的视频分析场景。实验表明,VideoAutoArena能有效区分顶尖LMM,揭示其优劣势。为进一步优化,我们推出辅助基准VideoAutoBench,由人工标注部分对决胜者,再以GPT-4o比对响应,形成低成本、可扩展的用户导向视频分析评估体系。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) with advanced video analysis capabilities have recently garnered significant attention. However, most evaluations rely on traditional methods like multiple-choice questions in benchmarks such as VideoMME and LongVideoBench, which are prone to lack the depth needed to capture the complex demands of real-world users. To address this limitation-and due to the prohibitive cost and slow pace of human annotation for video tasks-we introduce VideoAutoArena, an arena-style benchmark inspired by LMSYS Chatbot Arena's framework, designed to automatically assess LMMs' video analysis abilities. VideoAutoArena utilizes user simulation to generate open-ended, adaptive questions that rigorously assess model performance in video understanding. The benchmark features an automated, scalable evaluation framework, incorporating a modified ELO Rating System for fair and continuous comparisons across multiple LMMs. To validate our automated judging system, we construct a 'gold standard' using a carefully curated subset of human annotations, demonstrating that our arena strongly aligns with human judgment while maintaining scalability. Additionally, we introduce a fault-driven evolution strategy, progressively increasing question complexity to push models toward handling more challenging video analysis scenarios. Experimental results demonstrate that VideoAutoArena effectively differentiates among state-of-the-art LMMs, providing insights into model strengths and areas for improvement. To further streamline our evaluation, we introduce VideoAutoBench as an auxiliary benchmark, where human annotators label winners in a subset of VideoAutoArena battles. We use GPT-4o as a judge to compare responses against these human-validated answers. Together, VideoAutoArena and VideoAutoBench offer a cost-effective, and scalable framework for evaluating LMMs in user-centric video analysis.

视频分析多模态模型自动化评测用户模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。