arXiv:2601.22496cs.LGcs.AI2026-01被引 1

提出动作充分性准则,改进目标表示以提升长程强化学习控制效果。

Action-Sufficient Goal Representations

  • 引入信息论条件‘动作充分性’,确保目标表示能支持最优动作预测。
  • 实验证明,基于策略训练的目标表示比价值函数学习更优。
  • 适合研究长程决策、目标导向强化学习的学者参考。

在离线目标条件强化学习中,层次化方法将长时序任务分解为高层子目标预测与低层动作执行。关键设计在于目标表示——作为两层间接口的压缩目标编码。现有方法从价值学习中推导该表示,隐含假设:对价值估计足够信息也足以支持最优动作预测。我们证明,即使价值估计精确,此类表示仍可能使需不同最优动作的目标发生坍塌。为此,我们提出动作充分性,即保证最优动作预测所需的信息理论条件。证明价值充分性不蕴含动作充分性,并在离散环境中实证表明后者与控制成功更强相关。进一步发现,标准对数似然训练下低层策略自然诱导出近似动作充分的目标表示。实验显示,此类表示始终优于基于价值函数估计学习的表示。

原文摘要 · Abstract (English)

In offline goal-conditioned reinforcement learning (GCRL), hierarchical approaches decompose long-horizon tasks into high-level subgoal prediction and low-level action execution. A critical design choice in such architectures is the goal representation-the compressed encoding of goals that serves as the interface between these levels. Existing methods derive this representation from value learning, implicitly assuming that information sufficient for value estimation is adequate for optimal action prediction. We show that this assumption can fail even under exact value estimation, as such representations may collapse goals requiring distinct optimal actions. To address this, we introduce action sufficiency, an information-theoretic condition on goal representations necessary for optimal action prediction. We prove that value sufficiency, the preservation of sufficient information for value estimation, does not imply action sufficiency and empirically verify that the latter is more strongly associated with control success in a discrete environment. We further demonstrate that an actor-based representation, naturally induced by standard log-likelihood training of the low-level policy, is approximately action-sufficient. Empirically, our actor-based representations consistently outperform representations learned via value function estimation.

强化学习目标表示动作充分性层次决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。