arXiv:2409.17355cs.LG2024-09被引 1

从行为示范中学习智能体的风险偏好,提升逆强化学习的实用性。

Learning Utilities from Demonstrations in Markov Decision Processes

  • 用效用函数显式建模智能体的风险态度,突破传统风险中性假设。
  • 提出两种可证明高效的算法,在有限数据下学习风险偏好并分析样本复杂度。
  • 适用于需理解人类风险行为的场景,如自动驾驶与医疗决策。

本研究旨在从序列决策问题中的行为示范中提取有用知识。尽管人类在面对不确定性时普遍存在风险敏感行为,但大多数逆强化学习(IRL)模型假设智能体为风险中性。这种模型误设不仅影响准确性,还无法直接捕捉观测智能体的风险态度,而后者在许多应用中至关重要。本文提出一种新的马尔可夫决策过程(MDP)行为模型,通过效用函数显式表示智能体的风险态度。我们定义了效用学习(UL)问题:从MDP中的示范数据中推断智能体的风险态度,即其效用函数,并分析了该效用的局部可识别性。此外,我们设计了两种在有限数据条件下可证明高效的算法,并分析了它们的样本复杂度。最后,通过概念验证实验,我们实证验证了模型和算法的有效性。

原文摘要 · Abstract (English)

Our goal is to extract useful knowledge from demonstrations of behavior in sequential decision-making problems. Although it is well-known that humans commonly engage in risk-sensitive behaviors in the presence of stochasticity, most Inverse Reinforcement Learning (IRL) models assume a risk-neutral agent. Beyond introducing model misspecification, these models do not directly capture the risk attitude of the observed agent, which can be crucial in many applications. In this paper, we propose a novel model of behavior in Markov Decision Processes (MDPs) that explicitly represents the agent's risk attitude through a utility function. We then define the Utility Learning (UL) problem as the task of inferring the observed agent's risk attitude, encoded via a utility function, from demonstrations in MDPs, and we analyze the partial identifiability of the agent's utility. Furthermore, we devise two provably efficient algorithms for UL in a finite-data regime, and we analyze their sample complexity. We conclude with proof-of-concept experiments that empirically validate both our model and our algorithms.

逆强化学习风险偏好效用函数智能体建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。