arXiv:2607.21067cs.CL2026-07

用可解释的偏好矩阵提升文本生成个性化效果

PrefReward: Learning User Preference Matrix for Personalized Text Generation

论文配图:PrefReward: Learning User Preference Matrix for Personalized Text Generation
图 1 · 摘自论文原文
  • 构建用户风格偏好矩阵,显式建模个性特征
  • 在LongLaMP数据集上生成质量与可解释性双提升
  • 适合需要可控个性化生成的研究者与开发者

大型语言模型虽能通过用户历史和上下文生成个性化内容,但现有方法依赖模型参数中的隐式表示,难以解释用户偏好且难处理长上下文依赖。为此,我们提出PrefReward,一种新型偏好感知生成框架:第一阶段提取用户专属的偏好矩阵以总结其写作风格;第二阶段将该矩阵作为基于KL散度的奖励信号,融入解码过程。在LongLaMP数据集上的实验表明,PrefReward在生成质量与个性化可解释性方面均优于非个性化及基于检索的基线方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by leveraging user histories and contextual cues. However, most existing personalization approaches rely on implicit representations within model parameters, making it difficult to interpret user-specific preferences or effectively handle long-context dependencies. To address these challenges, we propose PrefReward, a novel preference-aware generative framework that explicitly models user styles through a structured preference matrix and integrates it into the decoding process as a reward signal. PrefReward consists of two stages: (1) extracting a user-specific preference matrix that summarizes individual stylistic tendencies, and (2) using the matrix to guide generation via a KL-divergence-based reward function. Experiments on the LongLaMP dataset show that PrefReward outperforms non-personalized and retrieval-based baselines in both generation quality and personalization interpretability.

个性化生成偏好建模可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。