用小模型引导大模型,实现高效个性化内容生成。
Unsupervised Human Preference Learning
- 用小型代理模型生成自然语言规则,指导大模型输出。
- 在邮件和文章数据集上显著优于基线方法。
- 无需微调大模型,适合资源有限的个人应用。
大型语言模型虽具备强大推理能力,却因缺乏用户偏好信息而难以生成个性化内容。现有方法如上下文学习和参数高效微调,在面对个体少量私有数据时仍显不足。本文提出一种新方法:利用小型参数模型作为偏好代理,生成自然语言规则来引导更大的预训练模型,实现高效个性化。该方法通过一个小型本地“方向盘”模型控制更大基础模型的输出,生成符合个人偏好的内容,同时充分利用大模型的知识与能力。关键在于无需对大模型进行微调。在邮件和文章数据集上的实验表明,该技术显著优于基线个性化方法。本方案以低数据与计算成本实现个性化,为语言模型的个性化应用开辟新路径。
原文摘要 · Abstract (English)
Large language models demonstrate impressive reasoning abilities but struggle to provide personalized content due to their lack of individual user preference information. Existing methods, such as in-context learning and parameter-efficient fine-tuning, fall short in capturing the complexity of human preferences, especially given the small, personal datasets individuals possess. In this paper, we propose a novel approach utilizing small parameter models as preference agents to generate natural language rules that guide a larger, pre-trained model, enabling efficient personalization. Our method involves a small, local "steering wheel" model that directs the outputs of a much larger foundation model, producing content tailored to an individual's preferences while leveraging the extensive knowledge and capabilities of the large model. Importantly, this personalization is achieved without the need to fine-tune the large model. Experimental results on email and article datasets, demonstrate that our technique significantly outperforms baseline personalization methods. By allowing foundation models to adapt to individual preferences in a data and compute-efficient manner, our approach paves the way for highly personalized language model applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。