提出自适应探索机制,提升连续时间LQ强化学习的效率与性能。
Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems
- 根据批评者和演员的方差动态调节熵正则化,实现自适应探索。
- 在无需调参情况下达到次线性后悔上界,性能媲美最优固定探索方法。
- 适合需要高效学习的连续控制场景,尤其适用于状态-控制耦合系统。
我们研究与文献\cite{huang2024sublinear}中相同类别的连续时间随机线性-二次(LQ)控制问题,其中波动率依赖于状态与控制,状态为标量值,且运行控制奖励缺失。提出一种无模型、数据驱动的自适应探索机制,通过批评者动态调整熵正则化,通过演员调整策略方差。不同于\cite{huang2024sublinear}中采用的恒定或确定性探索调度,后者需大量调参且忽略迭代中的学习进展,我们的方法在最小调参下显著提升学习效率。尽管具有灵活性,该方法仍实现与此前仅用固定探索调度获得的最佳已知无模型结果相当的次线性后悔上界。数值实验表明,自适应探索相比非自适应无模型与有模型方法,加速收敛并改善后悔表现。
原文摘要 · Abstract (English)
We study reinforcement learning (RL) for the same class of continuous-time stochastic linear--quadratic (LQ) control problems as in \cite{huang2024sublinear}, where volatilities depend on both states and controls while states are scalar-valued and running control rewards are absent. We propose a model-free, data-driven exploration mechanism that adaptively adjusts entropy regularization by the critic and policy variance by the actor. Unlike the constant or deterministic exploration schedules employed in \cite{huang2024sublinear}, which require extensive tuning for implementations and ignore learning progresses during iterations, our adaptive exploratory approach boosts learning efficiency with minimal tuning. Despite its flexibility, our method achieves a sublinear regret bound that matches the best-known model-free results for this class of LQ problems, which were previously derived only with fixed exploration schedules. Numerical experiments demonstrate that adaptive explorations accelerate convergence and improve regret performance compared to the non-adaptive model-free and model-based counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。