arXiv:2511.08594cs.CL2025-11ICLR被引 40

提出新方法提升大模型输出多样性,兼顾能力与对齐。

Diverse Preference Learning for Capabilities and Alignment

  • 解耦偏好学习中的熵与交叉熵项,精细控制生成多样性
  • 在难样本重复采样任务上准确率更高,输出更丰富多样
  • 能更好代表多元社会观点,适合需要包容性的应用

大型语言模型(LLMs)在社会中扮演越来越重要的角色,其能否体现多元视角至关重要。然而,近期研究表明,诸如强化学习人类反馈(RLHF)和直接偏好优化(DPO)等对齐算法显著降低了模型输出的多样性:对齐后的模型不仅文本结构和词汇选择趋于重复,解决问题的方式也更加趋同,反映的社会观点范围更窄。我们归因于偏好学习算法中使用的KL散度正则项,该机制系统性地高估多数意见,牺牲了输出多样性。为此,我们提出软偏好学习(Soft Preference Learning),将KL惩罚中的熵项与交叉熵项解耦,实现对生成多样性的细粒度控制。从能力角度看,使用该方法训练的模型在困难的重复采样任务上取得更高准确率,且输出具备更强的语义与词汇多样性;从对齐角度看,模型能表征更广泛的社会观点,并展现出改进的对数几率校准能力。值得注意的是,软偏好学习在效果上优于标准温度缩放,且构成帕累托改进。

原文摘要 · Abstract (English)

The ability of LLMs to represent diverse perspectives is critical as they increasingly impact society. However, recent studies reveal that alignment algorithms such as RLHF and DPO significantly reduce the diversity of LLM outputs. Not only do aligned LLMs generate text with repetitive structure and word choice, they also approach problems in more uniform ways, and their responses reflect a narrower range of societal perspectives. We attribute this problem to the KL divergence regularizer employed in preference learning algorithms. This causes the model to systematically overweight majority opinions and sacrifice diversity in its outputs. To address this, we propose Soft Preference Learning, which decouples the entropy and cross-entropy terms in the KL penalty - allowing for fine-grained control over LLM generation diversity. From a capabilities perspective, LLMs trained using Soft Preference Learning attain higher accuracy on difficult repeated sampling tasks and produce outputs with greater semantic and lexical diversity. From an alignment perspective, they are capable of representing a wider range of societal viewpoints and display improved logit calibration. Notably, Soft Preference Learning resembles, but is a Pareto improvement over, standard temperature scaling.

大模型对齐多样性偏好学习语义多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。