arXiv:2608.18008cs.LGcs.AI2026-08

用大模型反馈做奖励塑形,即使评分不准也能保持最优策略。

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

  • 将大模型进度评分作为有界势函数,构建目标增强的马尔可夫决策过程
  • 在四种配置下验证,即使大模型评分误差达基线20倍仍保最优策略
  • 适合研究大模型与强化学习融合的学者,尤其关注理论可靠性者

将大语言模型与强化学习结合日益流行,但其衍生奖励信号的理论基础常被忽略。本文将大语言模型规划器与强化学习控制器的混合架构形式化为一种目标增强的马尔可夫决策过程,并证明:当使用大模型的每状态进展评分作为有界势函数时,即便评分不准确,所生成的奖励塑形项仍能保持最优策略集合。该保证强于一般的大模型作为奖励的方法。我们在一个小规模MDP上通过四种势函数配置进行了数值验证,包括一个评分幅度是基础奖励20倍的对抗性场景。

原文摘要 · Abstract (English)

Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate. This guarantee is stronger than what general LLM-as-reward approaches provide. We verify the result numerically on a small MDP under four potential configurations, including an adversarial one scaled to twenty times the base reward magnitude.

强化学习大模型奖励塑形理论保障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。