arXiv:2509.22964cs.LGcs.AI2025-09

功能批评者让演员-评论家模型更稳定、探索更高效

Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration

  • 用策略条件价值函数解决评价策略动态变化问题
  • 理论证明在线性近似下算法收敛,无需完整覆盖假设
  • 首次建立功能批评者与高效探索的内在联系

演员-评论家(AC)框架在离线策略强化学习中表现优异,但面临评价策略持续变动的“移动目标”问题。功能批评者(即策略条件价值函数)通过显式将策略表示作为输入来缓解此问题。尽管概念上吸引人,以往方法仍难以超越标准AC。本文重新审视功能批评者,揭示其对稳定“致命三重奏”与“移动目标”间复杂互动至关重要。我们提出一个在线性功能近似下的收敛离线策略AC算法,克服了长期存在的理论与实践壁垒:采用基于目标的TD学习,支持动态行为策略,且无需“完全覆盖”假设。通过形式化双重信任-覆盖机制,为提升样本效率提供理论指导,严格控制行为策略更新与评论家重评估以最大化离线数据利用。其次,我们发现功能批评者与高效探索存在根本关联,现有无模型后验采样近似无法捕捉策略依赖不确定性,而功能批评者能弥补这一缺口。这是强化学习文献中的首个相关贡献。实践中,我们设计专用神经网络架构和极简AC算法,在DeepMind Control Suite上的初步实验中,不依赖标准实现技巧即达到与最先进方法相当的性能。

原文摘要 · Abstract (English)

The actor-critic (AC) framework has achieved strong empirical success in off-policy reinforcement learning but suffers from the "moving target" problem, where the evaluated policy changes continually. Functional critics, or policy-conditioned value functions, address this by explicitly including a representation of the policy as input. While conceptually appealing, previous efforts have struggled to remain competitive against standard AC. In this work, we revisit functional critics within the actor-critic framework and identify two critical aspects that render them a necessity rather than a luxury. First, we demonstrate their power in stabilizing the complex interplay between the "deadly triad" and the "moving target". We provide a convergent off-policy AC algorithm under linear functional approximation that dismantles several longstanding barriers between theory and practice: it utilizes target-based TD learning, accommodates dynamic behavior policies, and operates without the restrictive "full coverage" assumptions. By formalizing a dual trust-coverage mechanism, our framework provides principled guidelines for pursuing sample efficiency-rigorously governing behavior policy updates and critic re-evaluations to maximize off-policy data utility. Second, we uncover a foundational link between functional critics and efficient exploration. We demonstrate that existing model-free approximations of posterior sampling are limited in capturing policy-dependent uncertainty, a gap the functional critic formalism bridges. These results represent, to our knowledge, first-of-their-kind contributions to the RL literature. Practically, we propose a tailored neural network architecture and a minimalist AC algorithm. In preliminary experiments on the DeepMind Control Suite, this implementation achieves performance competitive with state-of-the-art methods without standard implementation heuristics.

强化学习演员评论家稳定训练高效探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。