评测长视频上下文下多模态智能体的网页任务能力
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
- 基于4小时手工视频教程构建2021个网页任务
- 人类在事实保留任务上达73.9%,模型仅13.3%
- 长视频任务中模型表现反而低于无视频输入
视频常用于获取文本和静态图像无法提供的任务信息。然而,现有智能体评测基准多忽略长上下文视频理解,聚焦文本或静态图像输入。为此,我们提出VideoWebArena(VideoWA),一个评估长上下文多模态智能体视频理解能力的基准。VideoWA包含2021个基于人工制作视频教程的网页任务,总时长约四小时。我们定义了两类长视频任务:技能保持与事实保持。技能保持任务评估智能体能否利用人类示范高效完成任务;事实保持任务则评估其从视频中检索指令相关知识的能力。结果显示,最优模型在事实保持任务上的成功率为13.3%,在事实问答对上为45.8%,远低于人类的73.9%和79.3%。在技能保持任务中,使用视频教程时长上下文模型表现更差,在WebArena任务中下降5%,在VisualWebArena任务中下降10.3%。本工作凸显了提升长上下文多模态模型智能体能力的必要性,并为未来长视频智能体发展提供测试平台。
原文摘要 · Abstract (English)
Videos are often used to learn or extract the necessary information to complete tasks in ways different than what text and static imagery alone can provide. However, many existing agent benchmarks neglect long-context video understanding, instead focusing on text or static image inputs. To bridge this gap, we introduce VideoWebArena (VideoWA), a benchmark for evaluating the capabilities of long-context multimodal agents for video understanding. VideoWA consists of 2,021 web agent tasks based on manually crafted video tutorials, which total almost four hours of content. For our benchmark, we define a taxonomy of long-context video-based agent tasks with two main areas of focus: skill retention and factual retention. While skill retention tasks evaluate whether an agent can use a given human demonstration to complete a task efficiently, the factual retention task evaluates whether an agent can retrieve instruction-relevant information from a video to complete a task. We find that the best model achieves 13.3% success on factual retention tasks and 45.8% on factual retention QA pairs, far below human performance at 73.9% and 79.3%, respectively. On skill retention tasks, long-context models perform worse with tutorials than without, exhibiting a 5% performance decrease in WebArena tasks and a 10.3% decrease in VisualWebArena tasks. Our work highlights the need to improve the agentic abilities of long-context multimodal models and provides a testbed for future development with long-context video agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。