用少量偏好数据快速定制大模型,让回复更贴合用户个性。
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
- 将个性化建模转为元学习问题,仅需少量标注偏好即可推断用户偏好函数。
- 构建超100万条合成偏好数据,实现对1500个虚拟用户的个性化生成。
- 在真人测试中70%回答符合用户偏好,适合需要高效个性化的应用场景。
有效的大模型个性化对虚拟助手、内容推荐等应用至关重要。受大模型上下文学习能力启发,我们提出少样本偏好优化(FSPO),将奖励建模重构为元学习问题。在FSPO中,大模型通过少量带标签的偏好数据快速推断出用户个性化奖励函数。同时引入用户描述合理化(RAT)机制,提升奖励建模与指令遵循能力,恢复使用理想用户描述时的性能。由于真实偏好数据难以大规模获取,我们精心设计合成偏好数据集的构建方式,利用公开大模型生成超过100万条个性化合成偏好。为实现从合成数据到真实用户的成功迁移,我们发现数据需兼具高多样性与内在一致性结构。我们在电影评论、教育和开放式问答三个领域对最多1500个合成用户进行了个性化开放生成评估,并开展受控人类实验。结果显示,FSPO在合成用户上的Alpaca Eval胜率高达87%,在真实人类用户开放式问答任务中达到70%胜率。
原文摘要 · Abstract (English)
Effective personalization of LLMs is critical for a broad range of user-interfacing applications such as virtual assistants and content curation. Inspired by the strong in-context capabilities of LLMs, we propose few-shot preference optimization (FSPO), an algorithm for LLM personalization that reframes reward modeling as a meta-learning problem. Under FSPO, an LLM learns to quickly infer a personalized reward function for a user via a few labeled preferences. FSPO also utilizes user description rationalization (RAT) to encourage better reward modeling and instruction following, recovering performance with the oracle user description. Since real-world preference data is challenging to collect at scale, we propose careful design choices to construct synthetic preference datasets for personalization, generating over 1M synthetic personalized preferences using publicly available LLMs. To successfully transfer from synthetic data to real users, we find it crucial for the data to exhibit both high diversity and coherent, self-consistent structure. We evaluate FSPO on personalized open-ended generation for up to 1,500 synthetic users across three domains: movie reviews, education, and open-ended question answering. We also run a controlled human study. Overall, FSPO achieves an 87% Alpaca Eval winrate in generating responses that are personalized to synthetic users and a 70% winrate with real human users in open-ended question answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。