arXiv:2508.02103cs.LGstat.ML2025-08被引 1

提出一种自适应连续时间强化学习方法,能随环境难度自动调整策略。

Instance-Dependent Continuous-Time Reinforcement Learning via Maximum Likelihood Estimation

  • 通过最大似然估计状态边际密度引导学习
  • 后悔界与奖励方差和观测精度相关,适配复杂度时可独立于测量策略
  • 引入随机采样调度提升样本效率,不增加成本

连续时间强化学习(CTRL)为动态环境中连续演化的交互提供了自然框架。尽管CTRL已展现出日益增长的实证成功,但其对不同问题难度的适应能力仍不清楚。本文研究了CTRL的实例依赖行为,提出一种基于最大似然估计(MLE)和通用函数近似的简单模型化算法。不同于直接估计系统动力学的方法,本方法估计状态边际密度以指导学习。我们建立了实例依赖的性能保证,推导出后悔界随总奖励方差和测量分辨率变化。值得注意的是,当观测频率随问题复杂度适当调整时,后悔界将不再依赖于特定测量策略。为进一步提升性能,算法引入随机测量调度,在不增加测量成本的前提下提高样本效率。这些结果揭示了设计能根据环境难度自动调整学习行为的CTRL算法的新方向。

原文摘要 · Abstract (English)

Continuous-time reinforcement learning (CTRL) provides a natural framework for sequential decision-making in dynamic environments where interactions evolve continuously over time. While CTRL has shown growing empirical success, its ability to adapt to varying levels of problem difficulty remains poorly understood. In this work, we investigate the instance-dependent behavior of CTRL and introduce a simple, model-based algorithm built on maximum likelihood estimation (MLE) with a general function approximator. Unlike existing approaches that estimate system dynamics directly, our method estimates the state marginal density to guide learning. We establish instance-dependent performance guarantees by deriving a regret bound that scales with the total reward variance and measurement resolution. Notably, the regret becomes independent of the specific measurement strategy when the observation frequency adapts appropriately to the problem's complexity. To further improve performance, our algorithm incorporates a randomized measurement schedule that enhances sample efficiency without increasing measurement cost. These results highlight a new direction for designing CTRL algorithms that automatically adjust their learning behavior based on the underlying difficulty of the environment.

强化学习连续时间自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。