Terra Nova打造多挑战融合的智能体测试环境,逼真模拟复杂决策场景。
Terra Nova: A Comprehensive Challenge Environment for Intelligent Agents
- 融合多种典型强化学习挑战于单一环境
- 需长期规划与跨变量协同理解才能通关
- 适合评估智能体深度推理能力的研究者
我们提出Terra Nova,一个受《文明5》启发的综合性挑战环境(CCE),用于强化学习研究。CCE指在单一环境中同时出现多个经典强化学习挑战(如部分可观测性、信用分配、表征学习、巨大动作空间等)。掌握该环境需在多个相互作用的变量间实现整合的长期理解。我们强调,该定义排除了将无关任务简单并行聚合的基准(如同时学会所有Atari游戏)。此类聚合多任务基准主要检验智能体能否记忆和切换无关策略,而非测试其在多重交互挑战中进行深层推理的能力。
原文摘要 · Abstract (English)
We introduce Terra Nova, a new comprehensive challenge environment (CCE) for reinforcement learning (RL) research inspired by Civilization V. A CCE is a single environment in which multiple canonical RL challenges (e.g., partial observability, credit assignment, representation learning, enormous action spaces, etc.) arise simultaneously. Mastery therefore demands integrated, long-horizon understanding across many interacting variables. We emphasize that this definition excludes challenges that only aggregate unrelated tasks in independent, parallel streams (e.g., learning to play all Atari games at once). These aggregated multitask benchmarks primarily asses whether an agent can catalog and switch among unrelated policies rather than test an agent's ability to perform deep reasoning across many interacting challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。