测试AI网页代理发现隐藏信息并正确推理的能力,发现其表现普遍不佳。
PATHWAYS: Evaluating Investigation and Context Discovery in AI Web Agents
- 设计250个多步任务,评估代理发现隐藏上下文的能力。
- 仅少数任务能获取关键证据,误导性表面信息导致准确率接近随机。
- 即使发现正确信息,也常无法整合到最终判断中,适合研究可信AI推理者。
我们提出PATHWAYS,一个包含250个多步决策任务的基准,用于测试基于网络的代理是否能够发现并正确使用隐藏的上下文信息。在封闭和开放模型中,代理通常能导航到相关页面,但在极少数情况下能获取决定性的隐藏证据。当任务需要推翻误导性的表面信号时,性能急剧下降至接近随机准确率。代理频繁虚构调查推理,声称依赖于从未访问过的证据。即使发现了正确的上下文,代理也常常无法将其整合进最终决策。提供更明确的指令虽能提升上下文发现能力,但往往降低整体准确率,揭示了程序合规性与有效判断之间的权衡。这些结果表明,当前网页代理架构缺乏可靠的自适应调查、证据整合和判断覆盖机制。
原文摘要 · Abstract (English)
We introduce PATHWAYS, a benchmark of 250 multi-step decision tasks that test whether web-based agents can discover and correctly use hidden contextual information. Across both closed and open models, agents typically navigate to relevant pages but retrieve decisive hidden evidence in only a small fraction of cases. When tasks require overturning misleading surface-level signals, performance drops sharply to near chance accuracy. Agents frequently hallucinate investigative reasoning by claiming to rely on evidence they never accessed. Even when correct context is discovered, agents often fail to integrate it into their final decision. Providing more explicit instructions improves context discovery but often reduces overall accuracy, revealing a tradeoff between procedural compliance and effective judgement. Together, these results show that current web agent architectures lack reliable mechanisms for adaptive investigation, evidence integration, and judgement override.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。