用监督学习从人类行为中推断奖励,更准确且适用范围广。
Supervised Reward Inference
- 基于行为-奖励配对数据,用监督学习直接建模行为到奖励的映射。
- 在低效行为数据上表现接近理论上限,元世界机器人任务中精度优异。
- 方法通用性强,可轻松扩展至动作和目标预测任务。
现有奖励推断方法通常假设人类示范遵循特定行为模型,但人类行为可能因规划或执行不佳而次优,或仅为传达目标而非达成目标。一种现有解决方案是在训练与部署分布一致的前提下,构建行为与已知奖励的配对数据集,并学习行为到奖励的映射;然而先前方法多受限于表格型场景。本文提出监督奖励推断(SRI),利用此类数据集,通过监督学习实现简洁而强大的推断。理论上,我们在标准假设下证明SRI渐近贝叶斯最优。实验表明,SRI在先前奖励推断基准上达到近天花板性能,在Meta-World机器人任务中,即使从任意次优示范中也能以允许的精度推断奖励。最后,我们展示了该框架在动作预测与目标预测任务中的自然泛化能力。
原文摘要 · Abstract (English)
Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models. However, humans often indicate their goals through a wide range of behaviors, from actions that are suboptimal due to poor planning or execution to behaviors intended to communicate goals rather than achieve them. One existing solution for inferring rewards from such behavior $\unicode{x2013}$ provided it is drawn from the same distribution at training and deployment $\unicode{x2013}$ is to construct a dataset of behavior paired with known rewards, and to learn the mapping from behavior to rewards; however, prior methods in this family face notable limitations, such as restrictions to tabular settings. Given such a dataset, we propose instead that supervised learning offers a parsimonious yet powerful solution, which we term Supervised Reward Inference (SRI). Theoretically, we prove that SRI is asymptotically Bayes-optimal under standard assumptions. Empirically, SRI achieves near-ceiling performance on a prior benchmark for reward inference from suboptimal behavior, while on Meta-World robotics tasks, it infers rewards from even arbitrarily suboptimal demonstrations as accurately as those demonstrations allow. Finally, we demonstrate our framework's universality with straightforward generalizations to action- and goal-prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。