arXiv:2510.18849cs.CLcs.AI2025-10被引 3

用批评修正策略让大模型更真实可控地匹配用户偏好。

Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning

  • 引入多维评分与文本批评的奖励模型,防止奖励漏洞。
  • 让模型自我修正输出,提升个性化学习效率。
  • 在长度控制下表现优于PPO,小模型胜过GPT-4.1。

将大语言模型忠实个性化以匹配个体用户偏好是一项关键但具挑战性的任务。尽管监督微调(SFT)快速达到性能瓶颈,标准的人类反馈强化学习(RLHF)也难以捕捉个性化细微差别。基于标量的奖励模型易出现奖励黑客问题,导致回复冗长且表面化。为此,我们提出批判-后编辑(Critique-Post-Edit)框架,实现更忠实、可控制的个性化。该框架包含两个核心组件:(1) 个性化生成式奖励模型(GRM),提供多维评分与文本批评,抵御奖励黑客;(2) 批判-后编辑机制,使策略模型根据批评自行修订输出,实现更精准高效的训练。在严格的长度控制评估中,该方法显著优于标准PPO。个性化Qwen2.5-7B平均胜率提升11%,个性化Qwen2.5-14B性能超越GPT-4.1。结果表明该方法为实现忠实、高效、可控的个性化提供了可行路径。

原文摘要 · Abstract (English)

Faithfully personalizing large language models (LLMs) to align with individual user preferences is a critical but challenging task. While supervised fine-tuning (SFT) quickly reaches a performance plateau, standard reinforcement learning from human feedback (RLHF) also struggles with the nuances of personalization. Scalar-based reward models are prone to reward hacking which leads to verbose and superficially personalized responses. To address these limitations, we propose Critique-Post-Edit, a robust reinforcement learning framework that enables more faithful and controllable personalization. Our framework integrates two key components: (1) a Personalized Generative Reward Model (GRM) that provides multi-dimensional scores and textual critiques to resist reward hacking, and (2) a Critique-Post-Edit mechanism where the policy model revises its own outputs based on these critiques for more targeted and efficient learning. Under a rigorous length-controlled evaluation, our method substantially outperforms standard PPO on personalization benchmarks. Personalized Qwen2.5-7B achieves an average 11\% win-rate improvement, and personalized Qwen2.5-14B model surpasses the performance of GPT-4.1. These results demonstrate a practical path to faithful, efficient, and controllable personalization.

个性化强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。