arXiv:2603.12180cs.CLcs.AI2026-03被引 1

对比智能体与人类在文档搜索中的策略性差异,发现智能体依赖盲目尝试而非有效规划。

Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

  • 构建2250个真实问题的MADQA基准,评估智能体在异构文档中的推理能力
  • 最佳智能体准确率接近人类,但需更多操作且仍落后于理想表现近20%
  • 揭示智能体缺乏战略规划,适合关注高效推理与评估方法的研究者

多模态智能体为自动化复杂文档密集型工作流提供了前景,但关键问题在于:它们展现的是真正策略性推理,还是仅靠随机试错?为此,我们提出MADQA,一个包含2250个由人类编写的问题、基于800份异构PDF文档的基准。依据经典测验理论设计,以最大化不同智能体能力水平间的区分度。为评估智能体行为,我们引入新评估协议,衡量准确性与努力成本之间的权衡。结果表明,尽管最优智能体在原始准确率上可媲美人类搜索者,但其成功的问题类型大不相同,且依赖暴力搜索弥补策略规划不足。它们未能缩小与理想(oracle)性能之间近20%的差距,持续陷入低效循环。我们公开数据集与评估工具包,助力推动从盲目检索转向精准、高效的推理范式。

原文摘要 · Abstract (English)

Multimodal agents offer a promising path to automating complex document-intensive workflows. Yet, a critical question remains: do these agents demonstrate genuine strategic reasoning, or merely stochastic trial-and-error search? To address this, we introduce MADQA, a benchmark of 2,250 human-authored questions grounded in 800 heterogeneous PDF documents. Guided by Classical Test Theory, we design it to maximize discriminative power across varying levels of agentic abilities. To evaluate agentic behaviour, we introduce a novel evaluation protocol measuring the accuracy-effort trade-off. Using this framework, we show that while the best agents can match human searchers in raw accuracy, they succeed on largely different questions and rely on brute-force search to compensate for weak strategic planning. They fail to close the nearly 20% gap to oracle performance, persisting in unproductive loops. We release the dataset and evaluation harness to help facilitate the transition from brute-force retrieval to calibrated, efficient reasoning.

智能体文档搜索评估基准策略推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。