arXiv:2607.02975cs.AIcs.CL2026-07

测试智能体在信息与权限分散时的表现,发现社交访问显著降低成功率。

Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve

  • 将任务拆分为直接、中介和角色隔离三种访问模式,模拟真实组织结构。
  • 所有模型在社交访问下成功率下降,最高仅0.65,较集中模式低0.11。
  • 适合关注智能体协作能力与信息获取障碍的研究者参考。

多数智能体评测将事实、工具与权限集中于单一接口,但真实组织中这些资源分布于不同人员。本文通过将18个客服任务转化为三种设置:直接访问、一个已知中间人,以及六个角色隔离的参与者(需自主发现能力),研究当任务与成功标准不变而访问方式改变时的影响。在4个模型、864次试验中,社交访问导致所有模型成功率下降;两个模型的置信区间不包含零。最新测试模型gpt-5.6-sol在社交访问下表现最佳,成功率为0.65,较集中间接访问下降0.11,其置信区间包含零。探索性比较显示,在社交访问下五组模型对之间存在差异,而集中式设置无一区分。事后按任务分块检验中,三对交互的p值在多重校正后仍显著。参考相对读者认为更宽差距源于无法获取必要信息。因数据无法区分评估代理无效请求与模拟参与者错误回复,标签仅反映轨迹终止位置,而非模型差异原因。

原文摘要 · Abstract (English)

Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. Incognita asks what happens when the task and success criterion stay fixed but access does not. We transform eighteen customer-service tasks into three settings: direct access, one known intermediary, and six role-isolated participants whose capabilities must be discovered. Across 864 trials with four models, social access reduced success for every model; the pre-specified intervals excluded zero for two. The latest tested model, gpt-5.6-sol, achieved the highest social-access success at 0.65, a 0.11 decrease from centralized indirect access with an interval that included zero. Exploratory comparisons separated five of six model pairs under social access, while neither centralized setting separated any at this sample size. In post-hoc task-blocked tests, three pairwise interaction $p$-values remained significant after multiplicity adjustment. A reference-relative reader associates the wider gaps with failures to obtain needed information. Because the data cannot distinguish ineffective requests by the evaluated agent from inaccurate replies by simulated participants, the reader's labels describe where trajectories stopped, not why models differed.

智能体评测信息隔离协作能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。