arXiv:2508.08509cs.CLcs.AI2025-08AAAI被引 11

让大模型根据用户偏好灵活调整回答,更公平地满足多样需求。

Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression

  • 用少量样本对比推理,动态适配不同用户偏好
  • 在两个新基准上优于主流对齐方法,表现更稳定
  • 适合需要个性化伦理决策的场景,如医疗或法律咨询

大型语言模型(LLMs)当前主要通过人类反馈强化学习(RLHF)进行对齐,但这类方法依赖标量奖励,仅能反映用户偏好的平均值。相比之下,多元对齐旨在捕捉多个属性上的多样化用户偏好,超越单纯追求有用性和无害性。为此,我们提出一种基于少样本对比回归的可调控多元对齐模型,能够根据个体用户偏好自适应调整。该方法利用上下文学习与推理,基于细粒度属性对多个回应选项进行比较并做出对齐选择。为评估算法,我们还通过改编道德完整性语料库(MIC)和HelpSteer2数据集构建了两个新的可调控多元对齐基准,分别验证其在价值对齐决策与奖励建模中的适用性。我们的少样本对比回归方法具有可解释性,兼容多种属性与不同大模型,且在多项基线及前沿方法中表现更优。本研究为多元对齐提供了新思路,推动了伦理AI的发展。

原文摘要 · Abstract (English)

Large language models (LLMs) are currently aligned using techniques such as reinforcement learning from human feedback (RLHF). However, these methods use scalar rewards that can only reflect user preferences on average. Pluralistic alignment instead seeks to capture diverse user preferences across a set of attributes, moving beyond just helpfulness and harmlessness. Toward this end, we propose a steerable pluralistic model based on few-shot comparative regression that can adapt to individual user preferences. Our approach leverages in-context learning and reasoning, grounded in a set of fine-grained attributes, to compare response options and make aligned choices. To evaluate our algorithm, we also propose two new steerable pluralistic benchmarks by adapting the Moral Integrity Corpus (MIC) and the HelpSteer2 datasets, demonstrating the applicability of our approach to value-aligned decision-making and reward modeling, respectively. Our few-shot comparative regression approach is interpretable and compatible with different attributes and LLMs, while outperforming multiple baseline and state-of-the-art methods. Our work provides new insights and research directions in pluralistic alignment, enabling a more fair and representative use of LLMs and advancing the state-of-the-art in ethical AI.

多元对齐大模型个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。