arXiv:2504.06659cs.LGcs.AI2025-04中稿 · ICML被引 3

用机器遗忘技术低成本提升大模型对齐效果

Leveraging Machine Unlearning for Cost-Efficient Preference Alignment

  • 通过双层优化量化删除负面样本对对齐的影响
  • 实验证明不同负面样本影响差异大,可精准筛选
  • 适合关注高效对齐与降低数据成本的研究者

尽管大语言模型(LLMs)的偏好对齐(PA)取得进展,主流方法如基于人类反馈的强化学习仍面临挑战:需要高质量的正向偏好数据集,获取成本高且计算开销大。机器遗忘技术可通过直接消除负面样本的影响提供替代方案。然而,现有研究多集中于经验验证,缺乏系统的定量分析。为此,我们提出一个将偏好对齐与机器遗忘关联的框架。通过双层优化,我们首先量化特定负面样本的遗忘对对齐性能的影响,发现其效果在不同样本间差异显著。基于此,我们提出核心问题:如何最优地选择和加权负面样本进行遗忘以最大化对齐效果?为此,我们设计了无监督对齐(U2A)方法,利用双层优化实现高效样本选择与遗忘。大量实验验证了该方法的有效性。代码已开源。

原文摘要 · Abstract (English)

Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human feedback face notable challenges. These approaches require high-quality datasets of positive preference examples, which are costly to obtain and computationally intensive. The LLM unlearning technique presents a promising alternative by directly removing the influence of negative examples. However, current research has primarily focused on empirical validation, lacking systematic quantitative analysis. To bridge this gap, we propose a framework linking PA with LLM unlearning. Through bi-level optimization, we first quantify how unlearning specific negative examples impacts PA performance. Our analysis reveals that these effects vary substantially across negative examples. Building on this insight, we pose a crucial question: how can we optimally select and weight negative examples for unlearning to maximize PA performance? To answer this, we propose Unlearning to Align (U2A), which leverages bi-level optimization to efficiently select and unlearn examples for optimal PA performance. We validate the proposed method through extensive experiments, with results confirming its effectiveness. Our code is available at https://github.com/muyiahhh/U2A.

大模型对齐机器遗忘高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。