arXiv:2606.06053cs.LG2026-06中稿 · RLC 2026

在模型不准确时仍能保证强化学习的稳定表现

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification

  • 采用带KL正则的在线学习框架,适应函数近似下的模型偏差
  • 首次在非可实现场景下给出高概率的KL后悔界,含显式偏差项
  • 适合研究鲁棒强化学习与实际应用中存在模型误差的场景

我们研究在一般函数近似和模型不准确条件下的KL正则化上下文老虎机与回合制强化学习。现有理论依赖可实现性假设,因此无法推广到模型不准确的情形,此时经典后悔界可能失效。本文提出适用于上下文老虎机与回合制强化学习的KL不准确建模形式,并分析基于回归的算法与吉布斯策略更新方法。建立了包含显式不准确项的高概率KL后悔界,当模型可实现时可退化为标准结果。

原文摘要 · Abstract (English)

We study KL-regularized contextual bandits and episodic reinforcement learning (RL) under general function approximation with model misspecification. Existing guarantees rely on realizability and therefore do not extend to misspecified models, where classical regret bounds may fail. This work introduces KL misspecification formulations for contextual bandits and episodic RL and analyzes regression-based algorithms with Gibbs policy updates. High-probability KL-regret guarantees with explicit misspecification terms are established, recovering the standard realizable KL-regularized setting as a special case.

强化学习不准确模型后悔界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。