解决奖励学习中部分可识别问题,提升决策鲁棒性。
Robustness in the Face of Partial Identifiability in Reward Learning
- 用最坏情况下的可行奖励集优化性能,增强鲁棒性
- 提出Rob-ReL算法,理论保障样本与迭代复杂度
- 适合需可靠奖励估计的强化学习应用
在奖励学习(ReL)中,我们只能获得关于未知目标奖励的反馈,目标是利用这些信息恢复该奖励以用于下游任务(如规划)。当反馈信息不足时,目标奖励仅部分可识别,即存在一个等可能的奖励集合(称为可行集),使得所有成员都与观测一致。此时,现有方法可能恢复出与真实奖励不同的函数,导致下游任务失败。本文提出一个通用框架,用于量化因可识别性问题导致的性能下降。基于此,我们设计一种稳健方法:在可行集中对最差情况下的奖励最大化性能。进一步,我们开发了针对策略偏好评估任务的Rob-ReL算法,并提供了其在样本复杂度和迭代复杂度上的理论保证。最后通过概念验证实验展示了该设置的有效性。
原文摘要 · Abstract (English)
In Reward Learning (ReL), we are given feedback on an unknown target reward, and the goal is to use this information to recover it in order to carry out some downstream application, e.g., planning. When the feedback is not informative enough, the target reward is only partially identifiable, i.e., there exists a set of rewards, called the feasible set, that are equally plausible candidates for the target reward. In these cases, the ReL algorithm might recover a reward function different from the target reward, possibly leading to a failure in the application. In this paper, we introduce a general ReL framework that permits to quantify the drop in "performance" suffered in the considered application because of identifiability issues. Building on this, we propose a robust approach to address the identifiability problem in a principled way, by maximizing the "performance" with respect to the worst-case reward in the feasible set. We then develop Rob-ReL, a ReL algorithm that applies this robust approach to the subset of ReL problems aimed at assessing a preference between two policies, and we provide theoretical guarantees on sample and iteration complexity for Rob-ReL. We conclude with a proof-of-concept experiment to illustrate the considered setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。