arXiv:2603.17631cs.LGcs.AI2026-03

提出首个可生成已知最优策略的随机强化学习基准框架。

Benchmarking Reinforcement Learning via Stochastic Converse Optimality: Generating Systems with Known Optimal Policies

  • 基于反向最优性构造带噪声的非线性系统,确保策略与价值函数最优。
  • 通过参数扰动和同伦变化自动生成多样化环境,支持算法对比。
  • 适合追求严谨评估的强化学习研究者使用,提升实验可复现性。

强化学习算法的客观比较极为复杂,其性能表现高度依赖环境设计、奖励结构以及算法与环境动态中的固有随机性。为应对这一挑战,我们通过将反向最优性扩展至具有噪声的离散时间、控制仿射、非线性系统,提出了一种严格的基准测试框架。该框架提供了在何种条件下预设的价值函数与策略对所构造系统为最优的充要条件,从而可通过同伦变化和随机参数系统地生成基准族。我们通过自动构建多种环境验证了该框架,展示了其在算法间进行受控且全面评估的能力。通过将标准方法与真实最优解对比,本工作为强化学习的精确与严谨基准测试提供了可复现的基础。

原文摘要 · Abstract (English)

The objective comparison of Reinforcement Learning (RL) algorithms is notoriously complex as outcomes and benchmarking of performances of different RL approaches are critically sensitive to environmental design, reward structures, and stochasticity inherent in both algorithmic learning and environmental dynamics. To manage this complexity, we introduce a rigorous benchmarking framework by extending converse optimality to discrete-time, control-affine, nonlinear systems with noise. Our framework provides necessary and sufficient conditions, under which a prescribed value function and policy are optimal for constructed systems, enabling the systematic generation of benchmark families via homotopy variations and randomized parameters. We validate it by automatically constructing diverse environments, demonstrating our framework's capacity for a controlled and comprehensive evaluation across algorithms. By assessing standard methods against a ground-truth optimum, our work delivers a reproducible foundation for precise and rigorous RL benchmarking.

强化学习基准测试最优性随机系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。