让大模型通过对话理解个人偏好并动态调整回应。
Aligning LLMs with Individual Preferences via Interaction
- 通过多轮对话隐式推断用户未言明的个性化偏好。
- 构建包含3000+轮对话的树状偏好数据集,支持动态对齐。
- 适合需要个性化交互的应用,如智能助手、客服系统。
随着大语言模型(LLMs)能力日益增强,将其行为与人类价值观和偏好对齐成为广泛应用的关键。以往研究聚焦于通用原则(如有用性、无害性、诚实性),却忽视了个体差异化的偏好,可能影响定制化体验。为此,我们训练具备“互动对齐”能力的LLMs,使其通过多轮对话隐式推断当前用户的未言明偏好,并动态调整后续行为与回应。方法上,首先构建由3,310个不同用户人格构成的多样化池,通过迭代自生成与过滤扩展;在不同人格引导下,利用多模型协作生成包含3,000+轮对话的树状偏好数据集;最后采用监督微调与强化学习进行优化。评估方面,建立ALOE(ALign With CustOmized PrEferences)基准,包含100个精心设计样本及配套指标,用于衡量对话中个性化对齐效果。实验表明,该方法能有效实现基于互动的动态个性化对齐。
原文摘要 · Abstract (English)
As large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption. While previous research focuses on general alignment to principles such as helpfulness, harmlessness, and honesty, the need to account for individual and diverse preferences has been largely overlooked, potentially undermining customized human experiences. To address this gap, we train LLMs that can ''interact to align'', essentially cultivating the meta-skill of LLMs to implicitly infer the unspoken personalized preferences of the current user through multi-turn conversations, and then dynamically align their following behaviors and responses to these inferred preferences. Our approach involves establishing a diverse pool of 3,310 distinct user personas by initially creating seed examples, which are then expanded through iterative self-generation and filtering. Guided by distinct user personas, we leverage multi-LLM collaboration to develop a multi-turn preference dataset containing 3K+ multi-turn conversations in tree structures. Finally, we apply supervised fine-tuning and reinforcement learning to enhance LLMs using this dataset. For evaluation, we establish the ALOE (ALign With CustOmized PrEferences) benchmark, consisting of 100 carefully selected examples and well-designed metrics to measure the customized alignment performance during conversations. Experimental results demonstrate the effectiveness of our method in enabling dynamic, personalized alignment via interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。