提出新方法防止模型更新后倒退,保持原有对齐能力。
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization
- 用约束优化检测并缓解更新后的性能倒退
- 在保持整体对齐效果的同时减少退化现象
- 适合关注模型长期稳定性的研究者
近期偏好对齐技术显著提升了大语言模型生成符合人类偏好和价值观文本的能力。然而,现有对齐指标通常只关注更新后的总体改进,忽略了关键问题:更新后对先前正确处理的数据出现回退(即回归)。这种现象可能源于在已良好对齐数据上的过度微调,导致过对齐和性能退化。为此,我们提出 FlipGuard,一种基于焦点注意力的约束优化方法,用于检测并缓解更新回归。具体而言,FlipGuard通过定制化的奖励表征识别性能下降,并在训练中施加约束,以促进与预对齐模型的条件一致性。全面实验表明,FlipGuard有效缓解了更新回归,同时保持优异的整体性能,并具备知识保留优势。
原文摘要 · Abstract (English)
Recent breakthroughs in preference alignment have significantly improved Large Language Models' ability to generate texts that align with human preferences and values. However, current alignment metrics typically emphasize the post-hoc overall improvement, while overlooking a critical aspect: regression, which refers to the backsliding on previously correctly-handled data after updates. This potential pitfall may arise from excessive fine-tuning on already well-aligned data, which subsequently leads to over-alignment and degeneration. To address this challenge, we propose FlipGuard, a constrained optimization approach to detect and mitigate update regression with focal attention. Specifically, FlipGuard identifies performance degradation using a customized reward characterization and strategically enforces a constraint to encourage conditional congruence with the pre-aligned model during training. Comprehensive experiments demonstrate that FlipGuard effectively alleviates update regression while demonstrating excellent overall performance, with the added benefit of knowledge preservation while aligning preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。