让大模型根据任务自动提炼用户偏好,节省上下文并提升个性化效果。
Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

- 用语言强化学习训练可复用的文本优化策略,动态生成任务相关偏好。
- 平均提升3.82分,在39组实验中33组表现更好,仅保留22.8%原始偏好令牌。
- 适合需要长期个性化、资源受限的智能体系统,尤其擅长避免无关信息干扰。
自然语言用户偏好为大语言模型个性化提供了可解释接口,但通用偏好摘要常包含与下游任务无关的信息。直接使用完整摘要会浪费上下文容量并引入跨任务干扰,而手动设计任务专属偏好视图又难以扩展。本文研究任务特定偏好适配:给定通用用户偏好摘要和下游任务,生成保留关键决策证据同时去除冗余信息的任务条件表示。为此提出训练无关的元学习框架AlignXada,通过语言强化学习迭代优化可复用的文本精炼策略。在13个任务与3个下游模型(共39个任务-模型组合)上的实验显示,AlignXada平均提升3.82分,改善33个组合,仅需原偏好令牌的22.8%,且在36个组合中优于RAG。扩展的忠实性分析表明,精炼后的偏好仍基本基于原始输入,同时保留任务相关的个性化信号,说明偏好侧适配是终身个性化智能体中通用记忆构建的有效补充。
原文摘要 · Abstract (English)
Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary therefore wastes context capacity and introduces cross-task distraction, while manually designing task-specific preference views is difficult to scale. In this work, we study \emph{task-specific preference adaptation}: given a universal user preference summary and a downstream task, derive a task-conditioned representation that preserves sufficient decision-relevant evidence while removing redundant context. To this end, we propose \textsc{AlignXada}, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones. The refinement policy is iteratively optimized by a meta learner through verbal reinforcement learning. Across 13 tasks and three downstream models (39 task--model cells), \textsc{AlignXada} achieves an average gain of 3.82 points, improving 33 cells while retaining only 22.8\% of the original profile tokens and outperforming RAG in 36 cells. An extended faithfulness analysis further shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。