量化大模型对齐中的价值权衡,揭示隐性副作用
Value Alignment Tax: Measuring Value Trade-offs in LLM Alignment
- 基于价值观理论构建评估框架,分析对齐干预引发的跨价值影响
- 发现对齐目标值时,非目标值常出现系统性偏移,且效果不均衡
- 适合关注模型伦理风险、对齐机制设计的研究者使用
现有价值对齐研究多静态描述价值关系,忽视提示、微调或偏好优化等干预手段如何重塑整体价值体系。实际中,对齐某一价值可能隐式改变其他价值,产生未被测量的价值权衡。本文提出VAT框架,通过衡量对齐带来的价值变化在关联价值间的传播程度,相对于目标收益的比率,量化价值权衡。该方法捕捉对齐干预下价值表达的系统级动态,可同时评估预期提升与意外副作用。基于以舒瓦茨价值观理论为基础的控制型情景-行为数据集,收集配对的前后规范判断,分析不同模型、价值与干预方式下的影响。结果表明,对齐常导致价值间的非均匀、结构性共变,暴露目标与非目标价值间的系统性权衡。这类效应在传统单一目标评估中不可见,但通过VAT可清晰呈现,凸显过程级对齐风险,并揭示大模型价值对齐的动态本质。数据与代码已开源。
原文摘要 · Abstract (English)
Existing work on value alignment typically characterizes value relations statically, ignoring how alignment interventions, such as prompting, fine-tuning, or preference optimization, reshape the broader value system. In practice, aligning a target value can implicitly shift other values, creating value trade-offs that remain largely unmeasured. We introduce VAT, a framework that quantifies value trade-offs by measuring how alignment-induced changes propagate across interconnected values relative to achieved on-target gain. VAT captures the system-level dynamics of value expression under alignment intervention, enabling evaluation of both intended improvements and unintended side effects. Using a controlled scenario-action dataset grounded in Schwartz value theory, we collect paired pre-post normative judgments and analyze alignment effects across models, values, and interventions. Results show that alignment often produces uneven and structured co-movement among values, revealing systematic trade-offs between target and non-target values. These effects are largely invisible under conventional target-only evaluation, but become evident via VAT, highlighting process-level alignment risks and offering new insights into the dynamic nature of value alignment in LLMs. Dataset and code are open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。