arXiv:2601.08173cs.AI2026-01ACL被引 1

测试智能体在真实工作场景中的持续学习与探索能力

The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios

  • 构建动态环境模拟新人代理在变化任务中学习
  • 现有模型在主动探索和持续学习上表现差
  • 适合评估智能体在真实部署中的可靠性

多模态大语言模型虽推动了工作流自动化,但现有研究多关注静态环境下的性能上限,忽视了真实世界中随机部署的鲁棒性。我们识别出三大挑战:动态任务调度、不确定下的主动探索、以及从经验中持续学习。为此,提出 method{},一个模拟“学徒”代理在新环境中持续探索的动态评估环境。不同于传统基准, method{} 从三方面评估代理:(1) 对优先级变化的流式任务进行上下文感知调度;(2) 通过主动探索减少幻觉的信息获取策略;(3) 从规则驱动与动态生成任务中提炼通用策略以实现持续演化。实验表明,顶尖智能体在动态环境中存在明显缺陷,尤其在主动探索和持续学习方面。本工作建立了一套评估代理可靠性的框架,推动评测从静态测试转向真实生产场景。

原文摘要 · Abstract (English)

The rapid evolution of Multi-modal Large Language Models (MLLMs) has advanced workflow automation; however, existing research mainly targets performance upper bounds in static environments, overlooking robustness for stochastic real-world deployment. We identify three key challenges: dynamic task scheduling, active exploration under uncertainty, and continuous learning from experience. To bridge this gap, we introduce \method{}, a dynamic evaluation environment that simulates a "trainee" agent continuously exploring a novel setting. Unlike traditional benchmarks, \method{} evaluates agents along three dimensions: (1) context-aware scheduling for streaming tasks with varying priorities; (2) prudent information acquisition to reduce hallucination via active exploration; and (3) continuous evolution by distilling generalized strategies from rule-based, dynamically generated tasks. Experiments show that cutting-edge agents have significant deficiencies in dynamic environments, especially in active exploration and continual learning. Our work establishes a framework for assessing agent reliability, shifting evaluation from static tests to realistic, production-oriented scenarios. Our codes are available at https://github.com/KnowledgeXLab/EvoEnv

智能体评估持续学习动态调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。