arXiv:2601.02986cs.CL2026-01ACL被引 3

让模型自动生成动态评估清单,更精准匹配用户偏好。

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist

  • 用可插拔清单生成器动态构建评估标准
  • 提升奖励模型准确率并增强下游生成效果
  • 适合个性化推荐与生成任务的场景

近期个性化奖励建模方法多依赖用户交互历史来对齐模型判断与个体偏好,但现有方法通常将用户上下文视为静态或隐式条件信号,难以捕捉人类判断的动态与多维特性。本文提出 P-Check,一种新型个性化奖励建模框架,通过训练一个即插即用的清单生成器,合成动态评估标准以指导奖励预测。为更好对齐个性化细微差别,引入基于判别能力的偏好对比准则加权策略,为各项标准分配显著性评分。大量实验表明,P-Check 不仅提升奖励准确性,还增强下游个性化生成性能,并在分布外(OOD)场景中保持鲁棒性。

原文摘要 · Abstract (English)

Recent approaches in personalized reward modeling have primarily focused on leveraging user interaction history to align model judgments with individual preferences. However, existing approaches largely treat user context as a static or implicit conditioning signal, failing to capture the dynamic and multi-faceted nature of human judgment. In this paper, we propose P-Check, a novel personalized reward modeling framework, designed to train a plug-and-play checklist generator that synthesizes dynamic evaluation criteria for guiding the reward prediction. To better align these checklists with personalized nuances, we introduce Preference-Contrastive Criterion Weighting, a training strategy that assigns saliency scores to criteria based on their discriminative power for personalized judgment. We conduct extensive experiments and demonstrate that P-Check not only improves reward accuracy but also enhances downstream personalized generation, and remains robust in OOD scenarios.

个性化奖励模型动态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。