arXiv:2503.07631cs.LGcs.CL2025-03被引 1

构建开放世界视觉问答新基准,揭示大模型工具使用与推理短板

OWLViz: An Open-World Benchmark for Visual Question Answering

  • 设计需融合视觉理解、网络搜索与专用工具的多能力任务
  • 顶尖模型Gemini 2.0仅达26.6%准确率,远低于人类69.2%
  • 适合关注多模态推理、智能体系统与真实场景应用的研究者

我们提出一个具有挑战性的开放世界视觉问答(OWLViz)基准。该任务要求回答简洁、明确的问题,需整合视觉理解、网络探索和专用工具使用等多种能力。人类在这些直观任务中达到69.2%的准确率,而当前最先进的视觉语言模型(VLMs)表现不佳,最佳模型Gemini 2.0仅达26.6%准确率。依赖有限视觉与视觉语言模型作为工具的现有智能体系统表现更差。这一显著性能差距揭示了多模态系统在工具选择与复杂推理链执行方面的重大局限,为推进实用人工智能研究指明了新方向。

原文摘要 · Abstract (English)

We present a challenging benchmark for the Open WorLd VISual question answering (OWLViz) task. OWLViz presents concise, unambiguous queries that require integrating multiple capabilities, including visual understanding, web exploration, and specialized tool usage. While humans achieve 69.2% accuracy on these intuitive tasks, even state-of-the-art VLMs struggle, with the best model, Gemini 2.0, achieving only 26.6% accuracy. Current agentic VLMs, which rely on limited vision and vision-language models as tools, perform even worse. This performance gap reveals significant limitations in multimodal systems' ability to select appropriate tools and execute complex reasoning sequences, establishing new directions for advancing practical AI research.

视觉问答多模态智能体基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。