arXiv:2605.13207cs.LG2026-05

提出切换继承度量,实现无需预设目标的层次化零样本强化学习。

Switching Successor Measures for Hierarchical Zero-shot Reinforcement Learning

论文配图:Switching Successor Measures for Hierarchical Zero-shot Reinforcement Learning
图 1 · 摘自论文原文
  • 基于前向-后向表示,从单一模型中同时学习高层子目标选择与低层控制策略。
  • 在目标导向和通用奖励任务上均超越非层次基线,媲美先进层次方法。
  • 突破固定时间抽象和人工设计子目标的限制,适用于更广泛的任务场景。

层次化强化学习通过将长时决策分解为简单子问题来提升泛化能力。然而,现有方法常依赖固定时间抽象或目标条件化目标,限制于目标到达类任务,难以适应通用奖励函数。本文提出切换继承度量,作为继承度量的扩展,在无需额外监督、固定时域或手动设计子目标的情况下,实现层次化零样本强化学习。我们证明切换继承度量可自然从经典继承度量中导出,并保持其结构特性。基于此,提出FB $π$-Switch算法,直接从前向-后向(FB)表示中提取高层子目标选择策略与低层控制策略,使层次行为从单一学习表示中涌现。在目标导向与通用奖励任务上的实验表明,FB $π$-Switch优于非层次基线,并在目标导向设置中达到顶尖层次方法水平。结果表明,结构化继承表示为超越目标到达任务的层次化零样本强化学习提供了灵活基础。

原文摘要 · Abstract (English)

Hierarchical reinforcement learning can improve generalization by decomposing long-horizon decision-making into simpler subproblems. However, existing approaches often rely on restrictive design choices, such as fixed temporal abstractions or goal-conditioned objectives, which largely confine them to goal-reaching tasks and limit their applicability to general reward functions. In this paper, we introduce switching successor measures, an extension of successor measures that enables hierarchical control in zero-shot reinforcement learning without additional supervision, fixed horizons, or manually designed subgoals. We show that switching successor measures arise naturally from classical successor measures while preserving their underlying structure. Building on this result, we propose FB $π$-Switch, an algorithm that extracts both a high-level subgoal-selection policy and a low-level control policy directly from forward-backward (FB) representations, allowing hierarchical behavior to emerge from a single learned representation. Experiments on both goal-conditioned and general reward-based tasks show that FB $π$-Switch improves over non-hierarchical baselines and matches state-of-the-art hierarchical methods in goal-conditioned settings. These results demonstrate that structured successor representations provide a flexible foundation for hierarchical zero-shot reinforcement learning beyond goal-reaching tasks. Our project website is available at: https://stestokth.github.io/switching-successors/.

强化学习层次化零样本继承度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。