构建开放长周期文本游戏,测试智能体实时持续学习能力。
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents

- 生成包含丰富实体与动态的世界,支持长期任务探索
- 发现顶级智能体仍远低于人类表现,存在巨大提升空间
- 验证短期记忆对各类智能体均有帮助,是关键训练组件
为使智能体在测试时持续学习,需具备有效探索、获取新知识与技能、保留情景记忆并进行长程规划的能力。为此,我们提出AgentOdyssey,一个可程序化生成具有丰富实体、世界动态与长周期任务的开放文本游戏评估框架。该框架突破传统机器学习中测试阶段不学习的假设,将智能体置于持续、长周期的交互环境中,实现学习与推理的交织部署。我们进一步设计多维度评估方法,不仅衡量游戏进展,还诊断世界知识获取、情景记忆、物体与动作探索、动作多样性及模型开销。在生成的游戏上评估多种智能体范式,实验结果揭示智能体关键能力的显著局限及其影响因素。尽管性能随基础模型增强而提升,当前最优智能体仍远低于人类水平,显示巨大改进空间。研究发现,短期记忆对多种智能体机制均具增益,是测试时训练的重要组成部分。
原文摘要 · Abstract (English)
For agents to learn continuously from interaction with the world at test time, they must be able to explore effectively, acquire new world knowledge and skills, retain relevant episodic experiences, and plan over long horizons. To evaluate these key abilities of test-time continual learning agents, we introduce AgentOdyssey, a novel evaluation framework that procedurally generates open-ended text games with rich entities, world dynamics, and long-horizon tasks. Critically, AgentOdyssey goes beyond the conventional machine learning assumption that learning does not occur at test time by placing agents in a continuous, long-horizon setting that interleaves learning and inference throughout deployment. We further propose a multifaceted evaluation methodology that measures not only game progress but also offers diagnostic tests on world knowledge acquisition, episodic memory, object and action exploration, action diversity, and model cost. We evaluate diverse agent paradigms in the generated games. Our experimental results reveal critical limits in agents' key abilities, as well as factors that influence their meaningful horizon. Although performance scales with stronger base models, even the top agent remains far below human performance, leaving substantial headroom for improvement. Among agent mechanisms, we find that short-term memory benefits multiple agent paradigms and is an important component of agent test-time training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。