arXiv:2607.13396cs.AIcs.CL2026-07中稿 · COLM

测试大模型在工具可靠性突变时的适应能力,发现顶尖模型更稳定。

Set-shifting Behavioral Test for Harnessed Agents

论文配图:Set-shifting Behavioral Test for Harnessed Agents
图 1 · 摘自论文原文
  • 用认知心理学的集转移实验设计,模拟工具可靠性悄然变化
  • 部分模型几轮后就固守旧工具,前沿模型仍持续调用可靠工具组
  • 适合研究大模型鲁棒性与行为可塑性的研究人员参考

当一个可靠的工具在任务进行中悄然改变时,大型语言模型代理会如何调整其工具选择?我们借鉴认知心理学中的集转移概念,研究代理对隐藏可靠性变化的适应能力。通过构建包含冗余工具和技能的环境,其中多个工具可解决同一任务但可靠性不同,采用分支调度策略,在环境中切换可靠工具组,并与稳定对照组对比,以分离每次变化对代理行为的影响。我们在一组配备约束装置的大型语言模型上开展实验,发现相同的变化引发不同模型表现出显著差异:一些模型在几轮内就固定使用某一工具,而另一些则持续尝试变化;低能力模型常忽略可靠工具组,而前沿模型则继续调用该组工具。我们提出一套量化指标来衡量可靠性变化后的代理行为。尽管策略提示能显著改变某些模型的行为,但结果凸显了代理在间接可观测上下文变化下的脆弱性。

原文摘要 · Abstract (English)

What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow the notion of set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our cognitive test for LLM agents mounts libraries of redundant tools and skills, in which many tools solve the same task but differ in hidden reliability. Using a branching schedule, we shift the reliable tool group in the environment and compare it with a stable control, allowing us to isolate the effect of each shift on the agent's behavior. We conduct our study on a panel of LLMs equipped with harnesses and show that the same set of shifts results in distinct behaviors across models: some latch onto a fixed routine within a few turns, whereas others continue to vary. Less capable models often omit the reliable tool group, while frontier models keep calling it alongside the other groups. We introduce a suite of measures to quantify agent behavior after reliability shifts. While policy prompting substantially alters behavior in some tested models, our findings highlight agents' brittleness when changes occur in indirectly observable context.

大模型行为工具选择鲁棒性测试认知类比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。