提出高效探索框架,让智能体更准更快推断隐藏约束。
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
- 基于可行代价集分析误差影响,设计动态优化策略
- 理论保证样本复杂度,实测在多环境中提升推理效率
- 适合需从专家行为逆向学习约束的强化学习研究者
在众多实际应用中,对目标函数施加约束的优化至关重要。然而,这些约束往往难以直接定义,需从专家智能体行为中推断,即逆向约束推断问题。逆向约束强化学习(ICRL)是解决该问题的常用方法,依赖交互环境中收集的训练样本。但现有采样策略的有效性和效率尚不明确。本文提出一种具有理论保障的高效探索框架,通过定义ICRL问题的可行代价集,分析状态转移和专家策略估计误差对推断约束可行性的影响。基于此,提出两种探索算法:一是动态减小代价估计的累积误差边界,二是策略性地将探索集中在可能最优的策略附近。两类算法均具备可计算的样本复杂度理论保证,且在多种环境中的实验验证了其有效性。
原文摘要 · Abstract (English)
Optimizing objective functions subject to constraints is fundamental in many real-world applications. However, these constraints are often not readily defined and must be inferred from expert agent behaviors, a problem known as Inverse Constraint Inference. Inverse Constrained Reinforcement Learning (ICRL) is a common solver for recovering feasible constraints in complex environments, relying on training samples collected from interactive environments. However, the efficacy and efficiency of current sampling strategies remain unclear. We propose a strategic exploration framework for sampling with guaranteed efficiency to bridge this gap. By defining the feasible cost set for ICRL problems, we analyze how estimation errors in transition dynamics and the expert policy influence the feasibility of inferred constraints. Based on this analysis, we introduce two exploratory algorithms to achieve efficient constraint inference via 1) dynamically reducing the bounded aggregate error of cost estimations or 2) strategically constraining the exploration policy around plausibly optimal ones. Both algorithms are theoretically grounded with tractable sample complexity, and their performance is validated empirically across various environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。