用逆优化设计分层决策,让底层策略更懂长期目标。
Hierarchical Decision Making with Structured Policies: A Principled Design via Inverse Optimization

- 从专家示范中逆推底层优化目标,保证策略与全局目标一致。
- 在资源分配和避障任务中,效率与决策质量均优于现有方法。
- 适合需要严格约束和长期规划的复杂控制场景。
分层决策框架对解决复杂控制任务至关重要,能将难题分解为可管理的子目标。然而,现有分层策略存在两大缺陷:(i) 基于强化学习的方法难以保证严格满足约束条件;(ii) 基于最优控制的方法常依赖短视且计算成本高昂的建模方式。为此,分层强化学习-最优控制架构成为有前景的范式。但该框架中底层优化的构建仍缺乏系统性设计,多依赖启发式或短视目标。本文提出一种原则性框架,将上层目标抽象与结构化底层决策有机结合。通过逆优化方法,从专家示范中推导底层问题结构,确保底层策略目标始终与整体长期任务目标对齐。我们在网络资源分配和连续避障两类任务上验证该方法,结果表明其在效率和决策质量上持续优于端到端强化学习、基于学习的最优控制及现有分层强化学习基线。
原文摘要 · Abstract (English)
Hierarchical decision-making frameworks are pivotal for addressing complex control tasks, enabling agents to decompose intricate problems into manageable subgoals. Despite their promise, existing hierarchical policies face critical limitations: (i) reinforcement learning (RL)-based methods struggle to guarantee strict constraint satisfaction, and (ii) optimal control (OC)-based approaches often rely on myopic and computationally prohibitive formulations. To reconcile these trade-offs, hierarchical RL-OC architectures have emerged as a promising paradigm. However, the formulation of the lower-level optimization within these frameworks remains underexplored, often relying on heuristic or myopic objectives. In this work, we propose a principled framework that systematically integrates upper-level goal abstraction with structured lower-level decision making. We adopt an inverse optimization approach to inform the structure of the lower-level problem from expert demonstrations, ensuring that the objective of the lower-level policy remains aligned with the overall long-term task goal. To validate the approach, our framework is evaluated on distinct decision making tasks: network-based resource allocation and continuous collision avoidance. Empirical results demonstrate that our method consistently outperforms strong baselines based on end-to-end RL, learning-augmented optimal control, and existing hierarchical RL approaches in both efficiency and decision quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。