通过对比用户偏好提升大模型个性化生成效果
Personalized LLM Decoding via Contrasting Personal Preference
- 在参数高效微调后,用奖励引导解码增强个性化
- 平均提升10.57%的ROUGE-L得分,无需外部奖励模型
- 适合需要快速个性化但无额外训练资源的场景
随着大语言模型在各类实际应用中的部署,个性化需求日益重要。尽管已有基于提示和训练的方法被广泛研究,但解码阶段的个性化算法仍被忽视。本文提出CoPe(Contrasting Personal Preference),一种在用户特定数据进行参数高效微调(PEFT)后使用的新型解码阶段方法。核心思想是通过最大化用户的隐式奖励信号,实现个性化的奖励引导解码。我们在五个开放式个性化文本生成任务上评估CoPe,结果表明其性能显著,平均提升10.57%的ROUGE-L得分,且无需外部奖励模型或额外训练过程。
原文摘要 · Abstract (English)
As large language models (LLMs) are progressively deployed in various real-world applications, personalization of LLMs has become increasingly important. While various approaches to LLM personalization such as prompt-based and training-based methods have been actively explored, the development of effective decoding-time algorithms remains largely overlooked, despite their demonstrated potential. In this paper, we propose CoPe (Contrasting Personal Preference), a novel decoding-time approach applied after performing parameter-efficient fine-tuning (PEFT) on user-specific data. Our core idea is to leverage reward-guided decoding specifically for personalization by maximizing each user's implicit reward signal. We evaluate CoPe across five open-ended personalized text generation tasks. Our empirical results demonstrate that CoPe achieves strong performance, improving personalization by an average of 10.57% in ROUGE-L, without relying on external reward models or additional training procedures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。