构建多网页问答基准,测试模型跨页面推理能力。
WebQuest: A Benchmark for Multimodal QA on Web Page Sequences
- 设计多页网页问答数据集,需跨页面整合信息。
- 主流模型在多页推理上表现明显落后于单页任务。
- 通过思维链提示提升模型多屏推理能力,适合评估网页代理性能。
随着强大多模态大模型的发展,构建具备更高自主性的网络代理以辅助用户在各类人机界面中检索信息和完成任务变得愈发可行。因此,亟需建立涵盖多种真实使用场景的挑战性基准。本文提出 WebQuest,一个需要跨多个相关网页进行推理的多页面问答数据集。与现有聚焦多步网页导航和任务完成的界面基准不同,本数据集评估信息提取、多模态检索以及从多页中组合信息的能力。WebQuest 包含三类问题:单屏问答、多屏问答和基于导航轨迹的问答。我们评估了 GPT-4V、Gemini Flash、Claude 3 等主流商用模型,以及 InstructBLIP、PaliGemma 等开源模型,发现单屏与多屏推理之间存在显著差距。最后,我们研究了思维链提示等推理时技术对提升多屏推理能力的效果。
原文摘要 · Abstract (English)
The rise of powerful multimodal LLMs has enhanced the viability of building web agents which can, with increasing levels of autonomy, assist users to retrieve information and complete tasks on various human-computer interfaces. It is hence necessary to build challenging benchmarks that span a wide-variety of use cases reflecting real-world usage. In this work, we present WebQuest, a multi-page question-answering dataset that requires reasoning across multiple related web pages. In contrast to existing UI benchmarks that focus on multi-step web navigation and task completion, our dataset evaluates information extraction, multimodal retrieval and composition of information from many web pages. WebQuest includes three question categories: single-screen QA, multi-screen QA, and QA based on navigation traces. We evaluate leading proprietary multimodal models like GPT-4V, Gemini Flash, Claude 3, and open source models like InstructBLIP, PaliGemma on our dataset, revealing a significant gap between single-screen and multi-screen reasoning. Finally, we investigate inference time techniques like Chain-of-Thought prompting to improve model capabilities on multi-screen reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。