arXiv:2512.12088cs.AIcs.LG2025-12被引 1

RPI让强化学习在架构和环境变化下仍保持稳定高效。

Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations

  • 通过恢复价值估计单调性,提升策略迭代稳定性
  • 在CartPole和倒立摆任务中早期达近优性能并持续保持
  • 适合对训练稳定性与超参敏感度要求高的实际应用

近期工作提出可靠策略迭代(RPI),在函数逼近设置下恢复了策略迭代的价值估计单调性。本文评估了RPI在经典控制任务CartPole和Inverted Pendulum上,面对神经网络结构与环境参数变化时的鲁棒性表现。相比DQN、Double DQN、DDPG、TD3和PPO,RPI能更早达到近优性能,并在训练过程中持续维持该策略。由于深度强化学习常受样本效率低、训练不稳及超参数敏感等问题困扰,这些结果凸显了RPI作为更可靠替代方法的潜力。

原文摘要 · Abstract (English)

In a recent work, we proposed Reliable Policy Iteration (RPI), that restores policy iteration's monotonicity-of-value-estimates property to the function approximation setting. Here, we assess the robustness of RPI's empirical performance on two classical control tasks -- CartPole and Inverted Pendulum -- under changes to neural network and environmental parameters. Relative to DQN, Double DQN, DDPG, TD3, and PPO, RPI reaches near-optimal performance early and sustains this policy as training proceeds. Because deep RL methods are often hampered by sample inefficiency, training instability, and hyperparameter sensitivity, our results highlight RPI's promise as a more reliable alternative.

强化学习策略迭代鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。