arXiv:2505.12350cs.LGcs.AI2025-05

通过统计保证组合策略,提升强化学习稳定性与性能

Multi-CALF: A Policy Combination Approach with Statistical Guarantees

  • 融合标准策略与理论保障策略,智能组合决策
  • 可证明收敛到目标集,且偏差与收敛时间有严格上限
  • 适合对可靠性要求高的控制任务应用

我们提出 Multi-CALF,一种基于相对价值提升智能组合强化学习策略的算法。该方法将标准RL策略与具备理论保障的备选策略相结合,既继承了形式化的稳定性保证,又通常优于单一策略。我们证明了组合策略以已知概率收敛至指定目标集,并给出了最大偏离和收敛时间的精确界限。在控制任务上的实证验证表明,该方法在保持稳定性的同时提升了性能。

原文摘要 · Abstract (English)

We introduce Multi-CALF, an algorithm that intelligently combines reinforcement learning policies based on their relative value improvements. Our approach integrates a standard RL policy with a theoretically-backed alternative policy, inheriting formal stability guarantees while often achieving better performance than either policy individually. We prove that our combined policy converges to a specified goal set with known probability and provide precise bounds on maximum deviation and convergence time. Empirical validation on control tasks demonstrates enhanced performance while maintaining stability guarantees.

强化学习策略组合稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。