arXiv:2602.12116cs.CL2026-02中稿 · ICLR被引 5

让大模型更懂个人偏好,测试时动态调整评分标准。

P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling

  • 用生成式方法构建可变的评分链和用户原型
  • 在多个基准上平均提升2.31%,对新用户泛化性强
  • 适合需要个性化响应的对话系统研发者

个性化大语言模型对齐旨在根据个体用户偏好调整输出,通常通过强化学习实现。核心挑战在于开放场景下获取准确的用户专属奖励信号。现有个性化奖励模型存在两大局限:(1) 将多样、场景特定的偏好简化为固定的小规模评估准则;(2) 在反馈有限的新用户上泛化能力差。为此,我们提出 P-GenRM——首个支持测试时用户级动态缩放的个性化生成式奖励模型。P-GenRM 将偏好信号转化为结构化评估链条,自适应生成不同场景下的角色设定与评分标准。它进一步将用户聚类为用户原型,并引入双粒度缩放机制:在个体层面,动态缩放并聚合每位用户的评分体系;在原型层面,融合相似用户偏好。该设计缓解了推断偏好的噪声,通过原型迁移增强对未见用户的泛化能力。实证结果表明,P-GenRM 在主流个性化奖励模型基准上达到最先进水平,平均提升 2.31%,并在分布外数据集上表现优异。值得注意的是,测试时用户级缩放带来额外 3% 提升,验证了其在测试阶段实现强个性化对齐的能力。

原文摘要 · Abstract (English)

Personalized alignment of large language models seeks to adapt responses to individual user preferences, typically via reinforcement learning. A key challenge is obtaining accurate, user-specific reward signals in open-ended scenarios. Existing personalized reward models face two persistent limitations: (1) oversimplifying diverse, scenario-specific preferences into a small, fixed set of evaluation principles, and (2) struggling with generalization to new users with limited feedback. To this end, we propose P-GenRM, the first Personalized Generative Reward Model with test-time user-based scaling. P-GenRM transforms preference signals into structured evaluation chains that derive adaptive personas and scoring rubrics across various scenarios. It further clusters users into User Prototypes and introduces a dual-granularity scaling mechanism: at the individual level, it adaptively scales and aggregates each user's scoring scheme; at the prototype level, it incorporates preferences from similar users. This design mitigates noise in inferred preferences and enhances generalization to unseen users through prototype-based transfer. Empirical results show that P-GenRM achieves state-of-the-art results on widely-used personalized reward model benchmarks, with an average improvement of 2.31%, and demonstrates strong generalization on an out-of-distribution dataset. Notably, Test-time User-based scaling provides an additional 3% boost, demonstrating stronger personalized alignment with test-time scalability.

个性化奖励模型大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。