arXiv:2604.13780cs.LGcs.AI2026-04被引 2

提出一种新型离线策略强化学习方法,支持任意行为策略下的高效信用分配。

Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces

  • 基于优势函数的软树备份机制,实现多步离线策略学习
  • 统一框架下支持任意行为策略,提升训练灵活性与效率
  • 适用于需高探索性的复杂强化学习任务,如机器人控制

软Q学习已成为一种通用的无模型熵正则化强化学习方法,通过在回报中加入对参考策略偏离的惩罚来优化。尽管取得成功,其多步扩展仍相对未被探索,且仅限于玻尔兹曼策略下的在线策略动作采样。本文首先给出软Q学习的形式化n步公式,随后引入一种新颖的软树备份算子,将该框架扩展至完全离线策略情形。最终,我们统一上述成果,提出Soft Q(λ),一个优雅的在线、离线策略、使用资格迹的框架,可在任意行为策略下实现高效的信用分配。推导结果提出了一种无模型学习熵正则化值函数的方法,可应用于未来的实验研究。

原文摘要 · Abstract (English)

Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a penalty on the divergence from a reference policy. Despite its success, the multi-step extensions of soft Q-learning remain relatively unexplored and limited to on-policy action sampling under the Boltzmann policy. In this brief research note, we first present a formal $n$-step formulation for soft Q-learning and then extend this framework to the fully off-policy case by introducing a novel Soft Tree Backup operator. Finally, we unify these developments into Soft $Q(λ)$, an elegant online, off-policy, eligibility trace framework that allows for efficient credit assignment under arbitrary behaviour policies. Our derivations propose a model-free method for learning entropy-regularised value functions that can be utilised in future empirical experiments.

强化学习离线策略资格迹熵正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。