arXiv:2605.08131cs.LG2026-05

让智能体主动与专家互动,从中学习其行为背后的奖励机制。

Interactive Inverse Reinforcement Learning of Interaction Scenarios via Bi-level Optimization

论文配图:Interactive Inverse Reinforcement Learning of Interaction Scenarios via Bi-level Optimization
图 1 · 摘自论文原文
  • 用双层优化框架建模互动中的逆强化学习问题
  • 算法在多场景实验中成功推断出专家奖励函数并生成有效交互策略
  • 适合研究人机协作、动态交互场景的智能体训练

逆强化学习(IRL)旨在从专家示范数据中学习最优奖励函数和对应策略。然而,传统IRL假设学习者与专家隔离,仅能被动观察,难以应用于需要主动交互的场景。为此,本文提出交互式逆强化学习(IIRL),使学习者在与专家互动过程中主动推断其奖励函数并制定交互策略。我们将IIRL建模为随机双层优化问题:下层优化学习解释专家行为的奖励函数,上层优化学习与专家交互的策略。为此设计双循环算法——双层交互场景逆强化学习(BISIRL),内层求解奖励函数,外层求解交互策略。我们形式化证明了BISIRL的收敛性,并通过大量实验验证了算法的有效性。

原文摘要 · Abstract (English)

Inverse reinforcement learning (IRL) learns a reward function and a corresponding policy that best fit the demonstration data of an expert. However, in the current IRL setting, the learner is isolated from the expert and can only passively observe the expert demonstrations. This limits the applicability of IRL to interactive settings, where the learner actively interacts with the expert and needs to infer the expert's reward function from the interactions. To bridge the gap, this paper studies interactive IRL (IIRL) where a learner aims to learn the reward function of an expert and a policy to interact with the expert during its interactions with the expert. We formulate IIRL as a stochastic bi-level optimization problem where the lower level learns a reward function to explain the behaviors of the expert, and the upper level learns a policy to interact with the expert. We develop a double-loop algorithm, Bi-level Interactive Scenarios Inverse Reinforcement Learning (BISIRL), which solves the lower-level problem in the inner loop and the upper-level problem in the outer loop. We formally guarantee that BISIRL converges and validate our algorithm through extensive experiments.

逆强化学习交互学习双层优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。