让大模型同时理解多种价值冲突,提升对齐效果
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
- 通过优化多值元指令,增强模型对多元价值的理解
- 在5个价值集合上实现8种价值的更好平衡,超越现有基线
- 无需微调,适用于黑盒与开源模型,适合多价值场景
上下文学习在不进行昂贵后训练的情况下,展示了将大语言模型(LLMs)与人类价值观对齐的巨大潜力,称为上下文对齐(ICA)。然而,模型对输入提示的理解仍保持中立,限制了其应对价值矛盾的能力——人类价值观本质上是多元的,常产生相互冲突的要求,例如刺激性与传统性。当前的ICA方法因此面临指令瓶颈问题,即模型难以在单个提示中协调多个目标价值,导致对齐不完整或存在偏差。为此,我们提出PICACO,一种新颖的多元式上下文对齐方法。无需微调,PICACO通过优化包含多重价值的元指令,以更好地激发模型对这些价值的理解并提升对齐效果。该方法通过最大化指定价值与模型输出之间的总相关性来实现,理论上强化了价值一致性并减少了干扰噪声,从而生成更有效的指令。在五个价值集上的大量实验表明,PICACO适用于黑盒与开源模型,优于若干近期强基线,在最多8种不同价值间实现了更好的平衡。
原文摘要 · Abstract (English)
In-Context Learning has shown great potential for aligning Large Language Models (LLMs) with human values, helping reduce harmful outputs and accommodate diverse preferences without costly post-training, known as In-Context Alignment (ICA). However, LLMs' comprehension of input prompts remains agnostic, limiting ICA's ability to address value tensions--human values are inherently pluralistic, often imposing conflicting demands, e.g., stimulation vs. tradition. Current ICA methods therefore face the Instruction Bottleneck challenge, where LLMs struggle to reconcile multiple intended values within a single prompt, leading to incomplete or biased alignment. To address this, we propose PICACO, a novel pluralistic ICA method. Without fine-tuning, PICACO optimizes a meta-instruction that incorporates multiple values to better elicit LLMs' understanding of them and improve alignment. This is achieved by maximizing the total correlation between specified values and LLM responses, which theoretically reinforces value conformity and reduces distractive noise, resulting in more effective instructions. Extensive experiments on five value sets show that PICACO works well with both black-box and open-source LLMs, outperforms several recent strong baselines, and achieves a better balance across up to 8 distinct values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。