arXiv:2608.14718cs.CVcs.CL2026-08

构建多轮工具协作视频理解基准,评估AI助手真实智能水平

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

论文配图:VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
图 1 · 摘自论文原文
  • 将视频理解变为多轮交互+外部工具调用的智能体任务
  • 前沿模型在该基准上准确率不足60%,远低于90%的旧标准
  • 适合评估下一代通用AI助手的推理与规划能力

视频理解是评估多模态大模型能力的基础任务。然而,现有领先模型在Video-MME榜单上已达到约90%准确率,表明单轮视频问答任务日益饱和,难以衡量先进多模态大模型的真实智能。为此,我们提出VideoGAIA,一个面向通用人工智能助手的智能体式视频理解基准。该基准超越传统单次问答,将视频理解重构为多轮、工具增强的交互过程,要求模型迭代感知视频、调用外部工具、获取补充信息,并跨轮次整合多模态证据。VideoGAIA包含271个由模型与人类共同设计的任务,覆盖多样且复杂的现实场景。每个视频-问题-答案实例均经三位人类专家独立验证,确保正确性与难度适中。所有评测模型,包括GPT-5.5和Kimi-K3等前沿模型,在VideoGAIA上准确率均低于60%,凸显其作为高质量、及时性基准对下一代多模态大模型评估的价值。我们希望VideoGAIA能推动视频理解从传统范式向智能体范式演进。

原文摘要 · Abstract (English)

Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.

视频理解智能体多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。