系统梳理深度强化学习中的奖励模型技术,助你理解如何让智能体真正听懂人类意图。
Reward Models in Deep Reinforcement Learning: A Survey
- 按数据来源、机制和学习范式分类,梳理奖励模型设计方法
- 总结评估奖励模型的指标与实践应用,覆盖主流场景
- 适合研究者快速掌握该领域脉络,尤其关注对齐人类目标的智能体设计
在强化学习中,智能体通过与环境持续交互并利用反馈来优化行为。为引导策略优化,引入奖励模型作为期望目标的代理,使智能体最大化累积奖励时也实现任务设计者的意图。近年来,学术界与工业界均高度关注能够紧密对齐真实目标并促进策略优化的奖励模型。本文对深度强化学习领域的奖励建模技术进行了全面综述。首先介绍奖励建模的背景与基础概念;随后按数据来源、机制与学习范式对近年方法进行分类概述;在此基础上,讨论这些技术的应用场景,并回顾奖励模型的评估方法;最后指出未来有前景的研究方向。本综述涵盖经典与新兴方法,填补了当前文献中对奖励模型系统性梳理的空白。
原文摘要 · Abstract (English)
In reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are introduced as proxies of the desired objectives, such that when the agent maximizes the accumulated reward, it also fulfills the task designer's intentions. Recently, significant attention from both academic and industrial researchers has focused on developing reward models that not only align closely with the true objectives but also facilitate policy optimization. In this survey, we provide a comprehensive review of reward modeling techniques within the deep RL literature. We begin by outlining the background and preliminaries in reward modeling. Next, we present an overview of recent reward modeling approaches, categorizing them based on the source, the mechanism, and the learning paradigm. Building on this understanding, we discuss various applications of these reward modeling techniques and review methods for evaluating reward models. Finally, we conclude by highlighting promising research directions in reward modeling. Altogether, this survey includes both established and emerging methods, filling the vacancy of a systematic review of reward models in current literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。