arXiv:2601.12078cs.CLcs.IR2026-01ACL被引 5

用强化学习优化用户画像,让大模型更懂你

Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization

  • 引入上下文相关强化学习,动态构建用户画像
  • 在9项任务中超越传统方法,生成质量更高
  • 适合需要个性化大模型的智能客服、推荐系统

大型语言模型在通用任务上表现优异,但适配个体用户仍具挑战。检索增强通过基于用户历史记录来调节模型输出,现有方法通常依赖语义相关性选择记录。我们指出相关性是不可靠的代理:某些记录虽语义相近,却因冗余或冲突信息反而降低生成质量。为此,我们提出PURPLE框架,通过上下文相关强化学习优化用户画像构建。与仅选最相关记录的贪婪策略不同,PURPLE将画像构建视为顺序敏感的生成过程,采用Plackett-Luce排序模型捕捉记录间的复杂依赖关系。利用参考响应似然提供的丰富语义反馈进行训练,使检索直接对齐生成质量。在九项个性化任务上的大量实验表明,PURPLE在效果和效率上均持续优于强基线方法,为用户画像优化提供了原则性强且可扩展的解决方案。

原文摘要 · Abstract (English)

Large language models (LLMs) excel at general-purpose tasks, yet adapting their responses to individual users remains challenging. Retrieval augmentation provides a lightweight alternative to fine-tuning by conditioning LLMs on user history records, and existing approaches typically select these records based on semantic relevance. We argue that relevance serves as an unreliable proxy for utility: a record may be semantically similar to a query yet fail to improve generation quality or even degrade it due to redundancy or conflicting information. To bridge this gap, we propose PURPLE, a contextual bandit framework that oPtimizes UseR Profiles for LLM pErsonalization. In contrast to a greedy selection of the most relevant records, PURPLE treats profile construction as an order-sensitive generation process and utilizes a Plackett-Luce ranking model to capture complex inter-record dependencies. By training with semantically rich feedback provided by the likelihood of the reference response, our method aligns retrieval directly with generation quality. Extensive experiments on nine personalization tasks demonstrate that PURPLE consistently outperforms strong heuristic and retrieval-augmented baselines in both effectiveness and efficiency, establishing a principled and scalable solution for optimizing user profiles.

大模型个性化强化学习检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。