arXiv:2503.01048cs.LG2025-03NAACL被引 15

用自生成数据快速个性化大模型,无需微调

Personalize Your LLM: Fake it then Align it

  • 自动生成用户偏好数据,避免依赖真实数据集
  • 在LaMP基准上平均性能提升40%,优于两种基线方法
  • 适合需要低成本快速适配的个性化场景

个性化大型语言模型对提升用户体验至关重要。现有方法多需为每位用户微调模型,成本高昂;基于检索的方法虽更高效,但依赖高质量数据集,难以普及。为此,我们提出CHAMELEON,一种可扩展且高效的个性化方法,利用(1)自生成的个人偏好数据和(2)表示编辑技术,实现快速、低成本的模型个性化。实验表明,在包括LaMP在内的多个任务上,CHAMELEON能有效适应个人偏好,提升指令微调模型表现,相较于两种基线方法在两种模型架构上平均提升40%。

原文摘要 · Abstract (English)

Personalizing large language models (LLMs) is essential for delivering tailored interactions that improve user experience. Many existing personalization methods require fine-tuning LLMs for each user, rendering them prohibitively expensive for widespread adoption. Although retrieval-based approaches offer a more compute-efficient alternative, they still depend on large, high-quality datasets that are not consistently available for all users. To address this challenge, we propose CHAMELEON, a scalable and efficient personalization approach that uses (1) self-generated personal preference data and (2) representation editing to enable quick and cost-effective personalization. Our experiments on various tasks, including those from the LaMP personalization benchmark, show that CHAMELEON efficiently adapts models to personal preferences, improving instruction-tuned models and outperforms two personalization baselines by an average of 40% across two model architectures.

大模型个性化自生成数据表示编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。