arXiv:2412.11155cs.LGcs.AI2024-12AAAI被引 4

揭示非指数折扣智能体在逆强化学习中的偏好不可辨识性

Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting

  • 针对非指数折扣(如双曲线折扣)的智能体,分析逆强化学习中奖励函数的可辨识性
  • 证明在非指数折扣下,仅凭行为无法唯一确定真实奖励函数
  • 提醒研究者:仅靠逆强化学习难以准确刻画此类智能体的真实偏好

逆强化学习(IRL)旨在通过观察智能体行为推断其偏好。通常将偏好建模为奖励函数 $R$,行为建模为策略 $π$。然而,多个不同的 $R$ 可能产生相同的 $π$,导致 $R$ 无法被完全确定,即存在部分不可辨识性。现有研究已对最优和Boltzmann理性智能体的不可辨识性进行了刻画,但均假设未来奖励以指数方式折现。这一假设在行为科学中受到挑战,因为人类更符合双曲线折现。本文首次针对非指数折现智能体(尤其是双曲线折现)建立了 IRL 中的偏好多重性理论。结果表明,在非指数折现情形下,通常无法获取足够信息来识别正确的最优策略,意味着 IRL 单独不足以充分刻画此类智能体的偏好。

原文摘要 · Abstract (English)

The aim of inverse reinforcement learning (IRL) is to infer an agent's preferences from observing their behaviour. Usually, preferences are modelled as a reward function, $R$, and behaviour is modelled as a policy, $π$. One of the central difficulties in IRL is that multiple preferences may lead to the same observed behaviour. That is, $R$ is typically underdetermined by $π$, which means that $R$ is only partially identifiable. Recent work has characterised the extent of this partial identifiability for different types of agents, including optimal and Boltzmann-rational agents. However, work so far has only considered agents that discount future reward exponentially: this is a serious limitation, especially given that extensive work in the behavioural sciences suggests that humans are better modelled as discounting hyperbolically. In this work, we newly characterise partial identifiability in IRL for agents with non-exponential discounting: our results are in particular relevant for hyperbolical discounting, but they also more generally apply to agents that use other types of (non-exponential) discounting. We significantly show that generally IRL is unable to infer enough information about $R$ to identify the correct optimal policy, which entails that IRL alone can be insufficient to adequately characterise the preferences of such agents.

逆强化学习偏好推断非指数折扣不可辨识性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。