arXiv:2602.12394cs.LG2026-02KDD被引 3

用合成数据提升大模型个性化,让提示词更懂用户偏好。

Synthetic Interaction Data for Scalable Personalization in Large Language Models

  • 构建动态用户偏好模拟系统,生成高保真多轮交互数据。
  • 在多个基准上实现任务性能、个性化与抗噪能力的全面提升。
  • 适合需要个性化提示优化但无法修改模型的研究者与开发者。

个性化提示为大语言模型(LLM)服务多样化用户提供了巨大潜力,但现有提示优化方法主要聚焦任务层面,忽视了用户的特定偏好和潜在约束。这一差距源于(i)缺乏高质量、隐私敏感的规模化个人化用户-大模型交互数据,以及(ii)个体偏好缺乏稳健的奖励信号。为此,我们提出高保真合成数据生成框架PersonaGym。不同于将个性化视为静态人格-偏好对的传统方法,PersonaGym通过代理式大模型系统建模动态偏好过程,模拟真实偏好行为与语义感知噪声,生成个性化多轮交互轨迹。基于此,我们发布PersonaAtlas——一个大规模、高质量、多样化的合成数据集,其多轮交互轨迹高度贴近真实世界中的偏好表达与噪声模式。我们进一步提出可扩展且模型无关的个性化提示优化框架PPOpt,该框架基于交互历史优化用户提示,无需修改部署的LLM。PPOpt采用“先推理后优化”范式,推断显式用户画像并据此条件化提示重写,以避免奖励劫持。其训练过程融合冷启动监督先验与结果驱动的多目标强化学习。大量实验表明,PPOpt在任务性能、个性化质量及对噪声与稀疏偏好信号的鲁棒性方面均优于当前最优基线。

原文摘要 · Abstract (English)

Personalized prompting offers large opportunities for deploying large language models (LLMs) to diverse users, yet existing prompt optimization methods primarily focus on task-level optimization while largely overlooking user-specific preferences and latent constraints of individual users. This gap is primarily due to (i) the absence of high-quality, privacy-sensitive data that capture personalized user-LLM interactions at scale, and (ii) the lack of robust reward signals for individual preferences. To overcome existing data limitations, we introduce a high-fidelity synthetic data generation framework called PersonaGym. Unlike prior work that treats personalization as static persona-preference pairs, PersonaGym models a dynamic preference process via an agentic LLM system to simulate realistic preference behaviors and semantic-aware noise in order to generate personalized multi-turn interaction trajectories. Using PersonaGym, we release PersonaAtlas, a large-scale, high-quality, and diverse synthetic dataset of high-fidelity multi-turn personalized interaction trajectories that closely mirror real-world preference expression and noise patterns. We further propose Personalized Prompt Optimization (PPOpt), a scalable and model-agnostic framework that optimizes user prompts based on interaction histories without modifying the deployed LLM. PPOpt adopts a reason-then-optimize paradigm that infers an explicit user profile and conditions prompt rewriting on the user profile to avoid reward hacking. Our training procedure for PPOpt integrates a cold-start supervised prior with outcome-driven multi-objective reinforcement learning. We present extensive experiments to demonstrate consistent improvements over state-of-the-art baselines in terms of task performance, personalization quality, and robustness to noisy as well as to sparse preference signals.

个性化提示优化合成数据大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。