arXiv:2510.02271cs.CLcs.AI2025-10被引 3

首个评估工具增强智能体多源信息检索能力的基准测试

InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents

  • 构建多源信息融合任务流水线,强制跨源依赖与真实场景约束
  • 实测14个顶尖智能体在医学金融等6大领域平均准确率仅38.2%
  • 揭示当前模型对工具选择与使用仍存在严重缺陷,22.4%失败源于误用

信息获取是人类的基本需求。然而现有大模型智能体过度依赖开放网络搜索,面临内容噪声大、不可靠的问题,且多数现实任务需精准的领域专有知识,难以通过网络获取。随着模型上下文协议(MCP)的出现,智能体可接入数千种专业工具,看似解决了此问题。但尚不清楚智能体能否有效利用这些工具,更关键的是能否将工具与通用搜索结合以解决复杂任务。为此,我们提出InfoMosaic-Bench,首个专注于工具增强型智能体多源信息获取的基准测试。涵盖医学、金融、地图、视频、网页及跨领域集成六大典型领域,任务要求结合通用搜索与领域专用工具。任务通过InfoMosaic-Flow生成,该流程基于验证过的工具输出设定任务条件,强制跨源依赖,并过滤掉可通过简单查表解决的捷径案例,确保任务可靠且非平凡。对14个前沿大模型智能体的实验揭示三大发现:(i) 仅依赖网络信息不足,GPT-5准确率仅38.2%,通过率67.5%;(ii) 领域工具带来选择性但不一致的增益,某些领域提升而另一些下降;(iii) 22.4%的失败源于错误的工具使用或选择,表明当前大模型在基础工具操作上仍有明显短板。

原文摘要 · Abstract (English)

Information seeking is a fundamental requirement for humans. However, existing LLM agents rely heavily on open-web search, which exposes two fundamental weaknesses: online content is noisy and unreliable, and many real-world tasks require precise, domain-specific knowledge unavailable from the web. The emergence of the Model Context Protocol (MCP) now allows agents to interface with thousands of specialized tools, seemingly resolving this limitation. Yet it remains unclear whether agents can effectively leverage such tools -- and more importantly, whether they can integrate them with general-purpose search to solve complex tasks. Therefore, we introduce InfoMosaic-Bench, the first benchmark dedicated to multi-source information seeking in tool-augmented agents. Covering six representative domains (medicine, finance, maps, video, web, and multi-domain integration), InfoMosaic-Bench requires agents to combine general-purpose search with domain-specific tools. Tasks are synthesized with InfoMosaic-Flow, a scalable pipeline that grounds task conditions in verified tool outputs, enforces cross-source dependencies, and filters out shortcut cases solvable by trivial lookup. This design guarantees both reliability and non-triviality. Experiments with 14 state-of-the-art LLM agents reveal three findings: (i) web information alone is insufficient, with GPT-5 achieving only 38.2% accuracy and 67.5% pass rate; (ii) domain tools provide selective but inconsistent benefits, improving some domains while degrading others; and (iii) 22.4% of failures arise from incorrect tool usage or selection, highlighting that current LLMs still struggle with even basic tool handling.

智能体评估多源信息工具使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。