arXiv:2508.03693cs.LG2025-08被引 1

提出新方法让机器人更快学懂人类偏好,还能保证学习结果可靠。

PAC Apprenticeship Learning with Bayesian Active Inverse Reinforcement Learning

  • 用信息论选择最有价值的示范场景,减少需要的人类演示次数。
  • 首次在有噪声示范下提供可证明的可靠性保障,确保学习效果不差。
  • 适合对安全要求高的场景,如自动驾驶和机器人控制。

随着人工智能系统自主性提升,使其决策与人类偏好一致变得至关重要。逆强化学习(IRL)通过示范数据推断人类偏好,并据此生成表现良好的代理策略。但在自动驾驶或机器人等高风险领域,仅具备良好平均性能不足,还需具备形式化保证的可靠策略——而获取足够多的人类示范以满足可靠性要求成本高昂。主动逆强化学习通过有策略地选择最具信息量的示范场景来缓解此问题。本文提出PAC-EIG,一种基于信息论的采集函数,直接针对代理策略的‘大概率近似正确’(PAC)保证,为带有噪声专家示范的主动IRL提供了首个理论保障。该方法最大化对代理策略遗憾度的信息增益,高效识别出需进一步示范的状态。此外,还提出了当奖励学习为主要目标时的Reward-EIG。针对有限状态-动作空间,本文给出了收敛性边界,揭示了先前启发式方法的失效模式,并通过实验验证了所提方法的优势。

原文摘要 · Abstract (English)

As AI systems become increasingly autonomous, reliably aligning their decision-making with human preferences is essential. Inverse reinforcement learning (IRL) offers a promising approach to infer preferences from demonstrations. These preferences can then be used to produce an apprentice policy that performs well on the demonstrated task. However, in domains like autonomous driving or robotics, where errors can have serious consequences, we need not just good average performance but reliable policies with formal guarantees -- yet obtaining sufficient human demonstrations for reliability guarantees can be costly. Active IRL addresses this challenge by strategically selecting the most informative scenarios for human demonstration. We introduce PAC-EIG, an information-theoretic acquisition function that directly targets probably-approximately-correct (PAC) guarantees for the learned policy -- providing the first such theoretical guarantee for active IRL with noisy expert demonstrations. Our method maximises information gain about the regret of the apprentice policy, efficiently identifying states requiring further demonstration. We also present Reward-EIG as an alternative when learning the reward itself is the primary objective. Focusing on finite state-action spaces, we prove convergence bounds, illustrate failure modes of prior heuristic methods, and demonstrate our method's advantages experimentally.

逆强化学习主动学习可靠性保障机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。