构建可精确诊断的强化学习测试环境,实现算法性能的透明评估。
Synthetic Monitoring Environments for Reinforcement Learning
- 设计可配置的连续控制任务,支持精确计算即时损失。
- 揭示状态/动作空间大小、奖励稀疏性等对算法性能的影响。
- 适合研究者用于系统分析算法在分布内/外的表现差异。
强化学习缺乏能实现精准、白盒诊断的基准测试环境。现有环境常混杂多种复杂因素且缺少真实最优策略指标,难以定位算法失败原因。本文提出合成监控环境(SMEs),一套无限规模的连续控制任务集合,具备完全可配置的任务特性与已知最优策略。SMEs支持精确计算瞬时遗憾,并通过严格的几何状态空间约束,实现分布内(WD)与分布外(OOD)的系统性评估。通过多维度消融实验对PPO、TD3和SAC进行验证,揭示了状态/动作空间规模、奖励稀疏性及最优策略复杂度等因素对算法在分布内外表现的影响。结果表明,SMEs为强化学习评估提供了标准化、透明化的测试平台,推动评价范式从经验性基准转向严谨科学分析。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) lacks benchmarks that enable precise, white-box diagnostics of agent behavior. Current environments often entangle complexity factors and lack ground-truth optimality metrics, making it difficult to isolate why algorithms fail. We introduce Synthetic Monitoring Environments (SMEs), an infinite suite of continuous control tasks. SMEs provide fully configurable task characteristics and known optimal policies. As such, SMEs allow for the exact calculation of instantaneous regret. Their rigorous geometric state space bounds allow for systematic within-distribution (WD) and out-of-distribution (OOD) evaluation. We demonstrate the framework's benefit through multidimensional ablations of PPO, TD3, and SAC, revealing how specific environmental properties - such as action or state space size, reward sparsity and complexity of the optimal policy - impact WD and OOD performance. We thereby show that SMEs offer a standardized, transparent testbed for transitioning RL evaluation from empirical benchmarking toward rigorous scientific analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。