arXiv:2410.04070cs.CLcs.AI2024-10ICLR被引 59

无需训练,在生成时动态调整大模型输出以匹配个人偏好。

PAD: Personalized Alignment of LLMs at Decoding-Time

  • 生成时引入个性化奖励机制,解耦生成与偏好
  • 在多个偏好下表现优于传统训练方法,且泛化能力强
  • 适合需要实时个性化响应的场景,如客服、创作助手

由于传统对齐方法存在计算成本高、数据需求大的问题,个性化偏好(因文化、教育、政治差异而异)的对齐成为挑战。本文提出解码时个性化对齐(PAD),一种在推理阶段实现个性化对齐的新框架,无需额外训练。通过独特的个性化奖励建模策略,该框架将文本生成过程与个性化偏好解耦,生成可泛化的逐标记奖励。PAD算法利用这些奖励引导解码,动态调整基础模型输出以适配用户偏好。大量实验表明,PAD不仅在对齐多样偏好方面优于现有训练型方法,还展现出对训练中未见偏好的显著泛化能力,并具备跨不同基础模型的可扩展性。该工作提升了大模型在实时应用中满足用户需求的能力,是个性化大模型对齐的重要进展。

原文摘要 · Abstract (English)

Aligning with personalized preferences, which vary significantly across cultural, educational, and political differences, poses a significant challenge due to the computational costs and data demands of traditional alignment methods. In response, this paper presents Personalized Alignment at Decoding-time (PAD), a novel framework designed to align LLM outputs with diverse personalized preferences during the inference phase, eliminating the need for additional training. By introducing a unique personalized reward modeling strategy, this framework decouples the text generation process from personalized preferences, facilitating the generation of generalizable token-level personalized rewards. The PAD algorithm leverages these rewards to guide the decoding process, dynamically tailoring the base model's predictions to personalized preferences. Extensive experimental results demonstrate that PAD not only outperforms existing training-based alignment methods in terms of aligning with diverse preferences but also shows significant generalizability to preferences unseen during training and scalability across different base models. This work advances the capability of LLMs to meet user needs in real-time applications, presenting a substantial step forward in personalized LLM alignment.

大模型对齐个性化生成解码优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。