arXiv:2504.20106cs.LGcs.AI2025-04中稿 · The 19th Conferenc…被引 5

用可动态组合的偏好向量,让大模型更灵活地平衡有用与安全。

Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors

  • 分离训练单一偏好模型,提取行为偏移作为向量
  • 测试时动态融合向量,实现无重训的多偏好调整
  • 用户可精细控制有用性与安全性的权衡,适合个性化部署

确保大语言模型既有用又无害是关键挑战:过度严格会导致过多拒绝,而过于宽松则可能生成有害内容。现有方法如基于人类反馈的强化学习(RLHF)和直接偏好优化(DPO)虽尝试平衡,但存在性能冲突、可控性差、扩展性不足等问题。为此,我们提出偏好向量(Preference Vector)框架,受任务算术启发。不将多个偏好合并于单一目标优化,而是分别训练各偏好模型,提取行为偏移作为偏好向量,并在测试时动态融合。该模块化方法支持细粒度、用户可控的偏好调节,且可无缝集成新偏好而无需重新训练。实验表明,该框架在不造成过度保守的前提下提升有用性,实现偏好权衡的平滑控制,并支持可扩展的多偏好对齐。

原文摘要 · Abstract (English)

Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessive refusals, while permissive models risk generating harmful content. Existing approaches, such as reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO), attempt to balance these trade-offs but suffer from performance conflicts, limited controllability, and poor extendability. To address these issues, we propose Preference Vector, a novel framework inspired by task arithmetic. Instead of optimizing multiple preferences within a single objective, we train separate models on individual preferences, extract behavior shifts as preference vectors, and dynamically merge them at test time. This modular approach enables fine-grained, user-controllable preference adjustments and facilitates seamless integration of new preferences without retraining. Experiments show that our proposed Preference Vector framework improves helpfulness without excessive conservatism, allows smooth control over preference trade-offs, and supports scalable multi-preference alignment.

大模型对齐偏好学习可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。