arXiv:2412.20834cs.CLcs.AI2024-12被引 6

分离偏好表示与生成,让模型更快适配个人喜好。

Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment

  • 将偏好表示与文本生成解耦,提升个性化对齐效率。
  • 新用户适配时间减少80%至90%,效果不逊于主流方法。
  • 适合需要快速响应个体偏好的对话系统应用。

将大型语言模型(LLMs)与通用人类偏好对齐已被证明对提升人机交互质量至关重要。然而,人类价值观在不同个体间存在固有差异,仅对齐通用偏好不足以满足需求。为此,根据个体反馈个性化模型成为有前景的解决方案。但该方法在对齐算法效率方面面临挑战。本文提出一种灵活的个体偏好对齐范式,通过将偏好表示与文本生成在模型中解耦,从根本上提升效率。我们在多个文本生成任务上验证了该方法,结果表明其生成质量与或优于基于参数高效微调(PEFT)的方法,同时使每个新个体偏好适配的额外训练时间减少80%至90%。

原文摘要 · Abstract (English)

Aligning Large Language Models (LLMs) with general human preferences has been proved crucial in improving the interaction quality between LLMs and human. However, human values are inherently diverse among different individuals, making it insufficient to align LLMs solely with general preferences. To address this, personalizing LLMs according to individual feedback emerges as a promising solution. Nonetheless, this approach presents challenges in terms of the efficiency of alignment algorithms. In this work, we introduce a flexible paradigm for individual preference alignment. Our method fundamentally improves efficiency by disentangling preference representation from text generation in LLMs. We validate our approach across multiple text generation tasks and demonstrate that it can produce aligned quality as well as or better than PEFT-based methods, while reducing additional training time for each new individual preference by $80\%$ to $90\%$ in comparison with them.

个性化效率优化偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。