让大模型同时符合多种人类价值观,解决价值冲突难题
Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation
- 通过降低不同价值观间的互信息,减少参数干扰
- 提出价值外推策略,高效探索最优价值权衡边界
- 适合需要多价值观平衡的AI安全与伦理场景
随着大语言模型(LLMs)的快速发展,将其对齐于人类价值观以确保安全与伦理已成为关键挑战。尤其在需同时考虑多个可能冲突的人类价值观时,这一问题更加突出。尽管已有多种对齐方法(如基于人类反馈的强化学习(RLHF)和直接偏好优化(DPO))被提出,但仍存在显著局限:1)多价值观优化中常不稳定且效率低下;2)难以有效处理价值观冲突。因此,现有方法通常难以实现多个价值观之间的最佳权衡。为应对该挑战,我们提出一种新框架——多值对齐(MVA)。该框架通过最小化不同人类价值观间的互信息,缓解因参数干扰导致的对齐退化。此外,我们提出一种价值外推策略,可高效探索帕累托前沿,从而构建一组具有多样化价值偏好的大模型。大量实验表明,MVA在对齐多个用户价值观方面始终优于现有基线方法。
原文摘要 · Abstract (English)
With the rapid advancement of large language models (LLMs), aligning them with human values for safety and ethics has become a critical challenge. This problem is especially challenging when multiple, potentially conflicting human values must be considered and balanced. Although several variants of existing alignment methods (such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO)) have been proposed to address multi-value alignment, they suffer from notable limitations: 1) they are often unstable and inefficient in multi-value optimization; and 2) they fail to effectively handle value conflicts. As a result, these approaches typically struggle to achieve optimal trade-offs when aligning multiple values. To address this challenge, we propose a novel framework called Multi-Value Alignment (MVA). It mitigates alignment degradation caused by parameter interference among diverse human values by minimizing their mutual information. Furthermore, we propose a value extrapolation strategy to efficiently explore the Pareto frontier, thereby constructing a set of LLMs with diverse value preferences. Extensive experiments demonstrate that MVA consistently outperforms existing baselines in aligning LLMs with multiple human values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。