用少量用户样本实现精准偏好建模,提升大模型个性化能力
PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- 基于一个或几个用户样例,通过推理机制识别个人偏好
- 在有限数据下仍保持高准确率和强泛化能力,媲美更大模型
- 适合需要个性化对齐的场景,如定制助手、内容推荐
奖励模型(RMs)是现有后训练方法的核心,通过在微调中提供反馈信号,使大语言模型输出与人类价值观对齐。然而,现有模型难以捕捉细微且用户特定的偏好,尤其在数据有限和跨领域情况下表现不佳。为此,我们提出PersRM-R1,首个基于推理的奖励建模框架,专门从仅一个或少数个人示例中识别并表示个人因素。为应对数据稀缺和泛化能力要求,该方法结合合成数据生成与两阶段训练流程:先监督微调,再强化学习微调。实验表明,PersRM-R1在相同规模模型中表现优于现有方法,并在准确性和泛化性上达到远大于自身规模模型的水平,为更有效的个性化大模型开辟了新路径。
原文摘要 · Abstract (English)
Reward models (RMs), which are central to existing post-training methods, aim to align LLM outputs with human values by providing feedback signals during fine-tuning. However, existing RMs struggle to capture nuanced, user-specific preferences, especially under limited data and across diverse domains. Thus, we introduce PersRM-R1, the first reasoning-based reward modeling framework specifically designed to identify and represent personal factors from only one or a few personal exemplars. To address challenges including limited data availability and the requirement for robust generalization, our approach combines synthetic data generation with a two-stage training pipeline consisting of supervised fine-tuning followed by reinforcement fine-tuning. Experimental results demonstrate that PersRM-R1 outperforms existing models of similar size and matches the performance of much larger models in both accuracy and generalizability, paving the way for more effective personalized LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。