通过信息论方法,主动选择最有价值的长轨迹来减少人类标注成本。
Toward Information Theoretic Active Inverse Reinforcement Learning
- 基于信息论设计新查询策略,一次收集多步轨迹而非单个动作。
- 在网格世界实验中显著降低所需人类示范数量,提升学习效率。
- 适合需要高效获取人类偏好的自主系统研究者参考。
随着人工智能系统日益自主,使其决策与人类偏好对齐变得至关重要。在自动驾驶或机器人等领域,难以手动编写表达这些偏好的奖励函数。逆强化学习(IRL)提供了一种从示范中推断未知奖励的可行方案。然而,获取人类示范成本高昂。主动逆强化学习(Active IRL)通过战略性地选择最具有信息量的场景来请求人类示范,从而减少所需的人类努力。以往工作通常仅允许逐状态询问一个动作,而本文提出并分析了收集更长轨迹的场景。我们提出了一个信息论驱动的采集函数,并设计了高效的近似算法,在一系列网格世界实验中验证其性能,为未来扩展到更一般场景奠定基础。
原文摘要 · Abstract (English)
As AI systems become increasingly autonomous, aligning their decision-making to human preferences is essential. In domains like autonomous driving or robotics, it is impossible to write down the reward function representing these preferences by hand. Inverse reinforcement learning (IRL) offers a promising approach to infer the unknown reward from demonstrations. However, obtaining human demonstrations can be costly. Active IRL addresses this challenge by strategically selecting the most informative scenarios for human demonstration, reducing the amount of required human effort. Where most prior work allowed querying the human for an action at one state at a time, we motivate and analyse scenarios where we collect longer trajectories. We provide an information-theoretic acquisition function, propose an efficient approximation scheme, and illustrate its performance through a set of gridworld experiments as groundwork for future work expanding to more general settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。