arXiv:2607.20536cs.AIcs.CL2026-07被引 1

构建复杂用户交互任务基准,测试智能体真实使用场景下的协作能力。

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

论文配图:AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
图 1 · 摘自论文原文
  • 基于模拟应用构建516个含模糊与约束的任务,强制智能体与用户多轮互动。
  • 顶尖模型Claude Opus 4.7在复杂任务上仅35.7%成功率,场景级指标更低至21.3%。
  • 适合研究人机协同、工具使用与真实交互的智能体开发者参考。

能够处理日常数字任务(如订购杂货)的工具使用智能体不仅需操作应用程序,还需与用户互动,例如提问澄清、请求确认或告知指令不可行。然而,现有评估基准未能涵盖此类互动的多样性,且多在小规模环境、有限非状态变更接口下运行。为弥补这一空白,我们提出AppWorld-UL——一个包含516个具有挑战性的任务的“用户在环”基准,要求多样化的智能体-用户交互。该基准建立在包含亚马逊、Spotify等9个流行模拟应用的AppWorld框架之上,通过系统性修改原任务引入歧义与约束,迫使智能体进行多种类型交互。用户行为由受控提示的LLM模拟,设定明确知识边界,相比先前无约束或过度刚性的方法更具可靠性。评估显示,当前最先进的大模型Claude Opus 4.7在AppWorld-UL上仅达48.6%成功率,复杂组合任务子集降至35.7%,在更严格的场景级指标下进一步下降至21.3%。分析表明,正确用户交互是成功的关键。该结果凸显了基准难度,也证明其推动用户在环工具使用智能体研究的潜力。

原文摘要 · Abstract (English)

Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user when the instruction is infeasible. However, current benchmarks for evaluating agent-user interactions do not capture the diversity of such interactions. Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a ``user-in-the-loop'' benchmark of 516 challenging tasks requiring diverse agent-user interactions. Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify, we systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction. User behavior is simulated by an LLM prompted to respond with carefully designed knowledge boundaries, offering more reliable simulation than the unconstrained or overly rigid alternatives used in prior work. Our evaluation reveals that a state-of-the-art LLM, Claude Opus 4.7, achieves only 48.6% success on AppWorld-UL, and only 35.7% on the harder, compositional subset. On the stricter, scenario-level metric, compositional task performance drops to only 21.3%. Our analysis reveals that correct user-interaction is crucial for success. This demonstrates the benchmark's difficulty and its potential to advance research on user-in-the-loop tool-use agents.

智能体交互工具使用用户在环基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。