arXiv:2504.11337cs.CL2025-04被引 6

通过数据筛选提升大模型多目标对齐,缓解有用性与安全性冲突

REWARD CONSISTENCY: Improving Multi-Objective Alignment from a Data-Centric Perspective

  • 基于奖励一致性筛选同时满足多目标的训练数据
  • 在优化安全性和有用性时,双方得分平均提升13.37%
  • 适合需要平衡多个评价维度的模型对齐研究者

语言模型的多目标偏好对齐常面临困境:优化某一人类偏好(如有用性)往往损害另一偏好(如无害性),因目标间存在内在冲突。现有工作多从算法角度解决,本文提出数据驱动的新思路,揭示能缓解此类冲突的数据类型。我们定义了奖励一致性(Reward Consistency, RC),识别出同时符合多个偏好目标的样本,从而降低训练中的冲突。通过梯度分析证明,符合RC的样本可自然抑制多目标优化过程中的性能退化。基于此,我们构建了奖励一致性采样框架,自动生成能有效缓解冲突的偏好数据集。所生成数据在同时优化无害性和有用性时,分别实现无害率与有用性胜率平均提升13.37%,且在不同多目标场景中均表现稳定。

原文摘要 · Abstract (English)

Multi-objective preference alignment in language models often encounters a challenging trade-off: optimizing for one human preference (e.g., helpfulness) frequently compromises others (e.g., harmlessness) due to the inherent conflicts between competing objectives. While prior work mainly focuses on algorithmic solutions, we explore a novel data-driven approach to uncover the types of data that can effectively mitigate these conflicts. Specifically, we propose the concept of Reward Consistency (RC), which identifies samples that align with multiple preference objectives, thereby reducing conflicts during training. Through gradient-based analysis, we demonstrate that RC-compliant samples inherently constrain performance degradation during multi-objective optimization. Building on these insights, we further develop Reward Consistency Sampling, a framework that automatically constructs preference datasets that effectively mitigate conflicts during multi-objective alignment. Our generated data achieves an average improvement of 13.37% in both the harmless rate and helpfulness win rate when optimizing harmlessness and helpfulness, and can consistently resolve conflicts in varying multi-objective scenarios.

多目标对齐数据筛选奖励一致性大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。