arXiv:2608.00155cs.AIcs.LG2026-08

测试大模型代理在持续任务流中的自进化能力,发现表现受场景和模型能力影响显著。

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

论文配图:AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
图 1 · 摘自论文原文
  • 构建任务流框架,模拟真实场景下的连续任务挑战。
  • 自进化效果随任务流复杂度变化,强模型未必持续受益。
  • 提醒研究者应以流式任务评估代理,而非孤立单任务。

大型语言模型代理可通过自身积累的经验持续自我进化。然而,现有研究多采用独立评估方式,导致其在真实流式环境中——即面对多样且复杂的任务流时的适应能力——仍不明确。为此,我们提出 AgentStream,一个统一评估框架,将代理基准组织为可配置的任务流,并在测试时实例化三种流式场景:孤立(Isolated)、顺序(Sequential)与交错(Interleaved),逐步改变任务流的范围与领域构成。在此框架下,我们组合评估了五种代表性自进化方法在三个前沿基础模型上的表现,分离出模型能力、方法架构与流式场景如何共同影响自进化。结果表明:自进化可靠性随流式场景变化;自进化收益受模型能力制约且非单调;无单一方法在所有模型与场景中占优。这些发现为跨模型与场景选择自进化方法提供了实证指导。总体而言,我们主张自进化代理应在真实任务流中评估,而非孤立的单任务设置。

原文摘要 · Abstract (English)

Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the \texttt{Isolated}, \texttt{Sequential}, and \texttt{Interleaved} streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution. Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and non-monotonic in model strength, and no single method dominates across models and scenarios. These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios. Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.

自进化流式任务大模型代理评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。