arXiv:2505.02666cs.CL2025-05综述被引 20

系统梳理大模型对齐中奖励设计的演进与方法

A Survey on Progress in LLM Alignment from the Perspective of Reward Design

  • 构建奖励建模的数学框架与实践路径
  • 揭示从单任务到多目标奖励机制的演进趋势
  • 为对齐研究提供理论框架与实践指导

奖励设计在将大语言模型与人类价值观对齐中起关键作用,是反馈信号与模型优化之间的桥梁。本文系统梳理了奖励建模的数学表述、构建实践及其与优化范式的关系,提出一个宏观层面的分类体系,从互补维度刻画奖励机制,为对齐研究提供概念清晰性与实践指导。大模型对齐的进展可理解为奖励设计策略的持续优化,近期发展体现出从基于强化学习到无强化学习优化、从单任务到多目标及复杂场景的范式转变。

原文摘要 · Abstract (English)

Reward design plays a pivotal role in aligning large language models (LLMs) with human values, serving as the bridge between feedback signals and model optimization. This survey provides a structured organization of reward modeling and addresses three key aspects: mathematical formulation, construction practices, and interaction with optimization paradigms. Building on this, it develops a macro-level taxonomy that characterizes reward mechanisms along complementary dimensions, thereby offering both conceptual clarity and practical guidance for alignment research. The progression of LLM alignment can be understood as a continuous refinement of reward design strategies, with recent developments highlighting paradigm shifts from reinforcement learning (RL)-based to RL-free optimization and from single-task to multi-objective and complex settings.

大模型对齐奖励设计强化学习综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。