arXiv:2605.17058cs.LG2026-05

提出多时标抽象框架,解决长程随机组合决策中的规划难题。

Learning Multi-Timescale Abstractions for Hierarchical Combinatorial Planning

论文配图:Learning Multi-Timescale Abstractions for Hierarchical Combinatorial Planning
图 1 · 摘自论文原文
  • 用潜空间树搜索与时变动作模型结合,实现自适应时间抽象。
  • 在多个基准上超越强基线,显著提升长程规划性能。
  • 适合资源受限下的复杂决策场景,如机器人路径规划。

序列随机组合优化(SSCO)因动作空间指数级膨胀、动态不确定性及资源有限的长期决策而极具挑战性。层次强化学习(HRL)虽提供自然分解方式,但高层策略需在半马尔可夫决策过程(SMDP)中处理时长可变的动作,导致难以构建适用于规划的世界模型。本文提出一种基于模型的层次化框架,直接应对该问题:将潜空间树搜索与面向SMDP的世界模型结合,通过多时标目标结构化潜空间动态,使状态转移幅度反映抽象动作的有效时间尺度,从而支持自适应时间抽象下的高效前瞻。此外,联合学习子目标条件化的预算策略,以实现上下文感知的资源分配。在多个挑战性的SSCO基准测试中,本方法优于现有强基线。

原文摘要 · Abstract (English)

The combination of exponentially large action spaces, stochastic dynamics, and long-horizon decision-making under limited resources makes Sequential Stochastic Combinatorial Optimization (SSCO) particularly challenging for reinforcement learning. Hierarchical Reinforcement Learning (HRL) offers a natural decomposition, but it places the high-level policy in a Semi-Markov Decision Process (SMDP) where actions have variable durations, making it difficult to learn a world model that is suitable for planning. We introduce a model-based hierarchical framework for sequential stochastic combinatorial decision-making that directly addresses this issue. Our method combines a latent-space tree-search planner with an SMDP-aware world model for variable-duration decisions. A multi-timescale objective structures the latent dynamics so that transition magnitudes reflect the effective temporal scales of abstract actions, enabling efficient lookahead under adaptive temporal abstraction. We further learn a subgoal-conditioned budget policy jointly with the world model to support context-aware resource allocation. Across challenging SSCO benchmarks, our method outperforms strong baselines.

层次强化学习组合优化时间抽象规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。