用自举回报建模人类选择,更好分离目标与信念,提升奖励学习鲁棒性。
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
- 用自举回报替代部分回报,融合对未来的估计来建模人类选择行为
- 即使人类信念错误,也能准确恢复对齐的奖励函数,而传统方法依赖正确信念
- 适合研究偏好学习、智能体对齐及人类行为建模的学者参考
随着人工智能代理生成越来越复杂的行为,手动编码人类偏好以引导这些代理变得愈发困难。为此,有研究建议代理从人类选择数据中学习偏好。这需要一个可解释选择数据的选择行为模型。对于部分状态与动作轨迹之间的选择,现有模型假设选择概率由部分回报或累积优势决定。本文提出一种基于自举回报的新模型,该模型在部分回报基础上增加对未来回报的估计。自举回报模型的优势源于其对人类信念的处理:与部分回报不同,它反映的是人类对环境的信念。此外,从累积优势推断奖励函数要求信念正确,而自举回报模型则无需此前提。为支持该模型,本文构建公理并证明了对齐定理,形式化表明此类模型能有效分离目标与信念,确保在自举回报基础上学习时恢复对齐奖励函数。该模型还具有更强鲁棒性:即使选择基于部分回报,通过自举回报模型仍能恢复对齐奖励;当选择基于累积优势且人机信念一致正确时亦成立。反之,若选择基于自举回报,使用部分回报或累积优势模型通常无法获得对齐奖励。
原文摘要 · Abstract (English)
As AI agents generate increasingly sophisticated behaviors, manually encoding human preferences to guide these agents becomes more challenging. To address this, it has been suggested that agents instead learn preferences from human choice data. This approach requires a model of choice behavior that the agent can use to interpret the data. For choices between partial trajectories of states and actions, previous models assume choice probabilities are determined by the partial return or the cumulative advantage. We consider an alternative model based instead on the bootstrapped return, which adds to the partial return an estimate of the future return. Benefits of the bootstrapped return model stem from its treatment of human beliefs. Unlike partial return, choices based on bootstrapped return reflect human beliefs about the environment. Further, while recovering the reward function from choices based on cumulative advantage requires that those beliefs are correct, doing so from choices based on bootstrapped return does not. To motivate the bootstrapped return model, we formulate axioms and prove an Alignment Theorem. This result formalizes how, for a general class of preferences, such models are able to disentangle goals from beliefs. This ensures recovery of an aligned reward function when learning from choices based on bootstrapped return. The bootstrapped return model also affords greater robustness to choice behavior. Even when choices are based on partial return, learning via a bootstrapped return model recovers an aligned reward function. The same holds with choices based on the cumulative advantage if the human and the agent both adhere to correct and consistent beliefs about the environment. On the other hand, if choices are based on bootstrapped return, learning via partial return or cumulative advantage models does not generally produce an aligned reward function.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。