测试大模型工具调用行为是否稳定可靠。
How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

- 设计实验重复调用同一智能体,记录工具选择顺序与参数。
- 发现相同任务下工具调用序列一致性不足,存在随机波动。
- 适合关注AI系统可靠性与部署风险的研究者和工程师。
具备工具调用能力的大语言模型代理正日益应用于生产系统,但一个基础性可靠性问题仍缺乏深入探讨:同一个智能体在两次相同调用中是否会表现出一致行为?本文对多步工具调用代理的行为一致性进行了系统的实证研究,衡量代理在重复调用相同任务时,是否以相同顺序选择相同的工具并传入相同的参数。不同于以往针对仅搜索或自由文本动作的ReAct型代理的一致性研究,本文聚焦于具有类型化参数和实际副作用的结构化工具调用接口这一更丰富的场景。
原文摘要 · Abstract (English)
Large language model (LLM) agents with tool-calling capabilities are increasingly deployed in production systems, yet a fundamental reliability question remains under-explored: does the same agent behave the same way twice? We present a systematic empirical study of behavioral consistency in multi-step tool-calling agents, measuring whether agents select the same tools, in the same order, with the same arguments, across repeated identical invocations. Unlike prior work on consistency in ReAct-style agents(search-only, free-text actions), we study the richer setting of structured tool-calling interfaces with typed parameters and consequential side effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。