arXiv:2504.02623cs.AI2025-04被引 7

用多任务动态切换测试LLM智能体的适应能力

Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions

  • 设计多任务关联测试场景,模拟真实复杂需求变化
  • 覆盖固定任务数下的所有任务切换模式,评估动态适应性
  • 适用于想提升工具调用鲁棒性的研究人员和开发者

大型语言模型(LLMs)凭借其强大的理解与规划能力,在工具调用代理方面展现出巨大潜力。用户越来越多地依赖基于LLM的代理通过迭代交互完成复杂任务。然而,现有基准大多仅在单任务场景下评估代理,难以反映真实世界的复杂性。为此,我们提出多任务工具基准(Multi-Mission Tool Bench)。该基准中每个测试用例包含多个相互关联的任务,要求代理动态适应不断变化的需求。此外,该基准探索了固定任务数量下的所有可能任务切换模式。我们还提出一种多代理数据生成框架构建基准,并设计了一种新的动态决策树方法,用于评估代理决策的准确性和效率。在多种开源与闭源LLM上的实验揭示了影响代理鲁棒性的关键因素,并为工具调用领域提供了可操作的洞见。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate strong potential as agents for tool invocation due to their advanced comprehension and planning capabilities. Users increasingly rely on LLM-based agents to solve complex missions through iterative interactions. However, existing benchmarks predominantly access agents in single-mission scenarios, failing to capture real-world complexity. To bridge this gap, we propose the Multi-Mission Tool Bench. In the benchmark, each test case comprises multiple interrelated missions. This design requires agents to dynamically adapt to evolving demands. Moreover, the proposed benchmark explores all possible mission-switching patterns within a fixed mission number. Specifically, we propose a multi-agent data generation framework to construct the benchmark. We also propose a novel method to evaluate the accuracy and efficiency of agent decisions with dynamic decision trees. Experiments on diverse open-source and closed-source LLMs reveal critical factors influencing agent robustness and provide actionable insights to the tool invocation society.

LLM代理多任务评估工具调用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。