让大模型学会适应动态变化的个人偏好,解决人与人之间价值冲突问题。
Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning

- 通过成对微调框架,让模型同时学习矛盾偏好。
- 在多选任务中准确率达96.6%,生成质量得分8.69,领先传统方法。
- 只需少量历史数据就能快速推断用户偏好,适合个性化应用。
大语言模型在对齐普遍人类偏好方面已取得显著进展,但在适应个体化、动态化的个人偏好方面仍面临挑战。本文提出一种新框架——偏好成对微调(Preference-Paired Fine-Tuning, PFT),旨在使模型能够处理相互冲突且持续演变的个体偏好。我们构建了一个新数据集——价值冲突困境(Value Conflict Dilemma, VCD),包含涉及矛盾人类偏好的情境,用于评估该方法。实验表明,PFT在多选分类任务中达到最高96.6%准确率,在开放式生成任务中获得8.69分的最优得分,显著优于DPO、SFT及部分传统训练方法,尤其在处理偏好冲突时表现突出。此外,在仅依赖有限用户历史数据的情况下,模型可快速推断偏好向量,相比单偏好模型,用户特定偏好对齐提升44.76%。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have significantly improved the alignment of models with general human preferences. However, a major challenge remains in adapting LLMs to individual preferences, which are not only diverse but also dynamic. In this paper, we introduce a novel framework, Preference-Paired Fine-Tuning (PFT), designed to align models with contradictory and evolving individual preferences. We present a new dataset, Value Conflict Dilemma (VCD), which includes scenarios that involve conflicting human preferences, facilitating the evaluation of our approach. Our experiments demonstrate that PFT outperforms single-preference training methods, achieving up to 96.6% accuracy in multi-choice classification tasks and the highest open-ended generation score of 8.69. PFT also shows significant improvements over DPO, SFT and some traditional training methods, especially when handling conflicting preferences. Additionally, with limited user history data, models can inferring preference vector rapidly, achieving a 44.76% improvement in user-specific preference alignment in comparison to single-preference models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。