用影响函数自动修正大模型偏差,无需人工干预。
Correcting Large Language Model Behavior via Influence Function
- 通过影响函数定位导致错误输出的训练数据
- 基于影响分布优化模型行为,修复不当输出
- 比依赖人工标注的方法更高效且可解释
近年来的AI对齐技术显著提升了大语言模型(LLMs)与静态人类偏好的一致性。然而,人类偏好的动态性使部分训练数据过时甚至错误,导致模型偏离当前社会规范。现有方法如持续对齐的数据筛选或人工修正旧数据,均需大量人力成本。为此,我们提出一种新方法——基于影响函数召回与后训练的大语言模型行为修正(LANCET),全程无需人工参与。LANCET包含两个阶段:(1) 利用影响函数识别显著影响模型不良输出的训练数据;(2) 采用影响函数驱动的Bregman优化(IBO)技术,依据这些影响分布调整模型行为。实验表明,LANCET能有效高效地纠正大模型的不当行为,性能优于依赖人类偏好收集的方法,并提升了模型学习人类偏好的可解释性。
原文摘要 · Abstract (English)
Recent advancements in AI alignment techniques have significantly improved the alignment of large language models (LLMs) with static human preferences. However, the dynamic nature of human preferences can render some prior training data outdated or even erroneous, ultimately causing LLMs to deviate from contemporary human preferences and societal norms. Existing methodologies, whether they involve the curation of new data for continual alignment or the manual correction of outdated data for re-alignment, demand costly human resources. To address this challenge, we propose a novel approach, Large Language Model Behavior Correction with Influence Function Recall and Post-Training (LANCET), which requires no human involvement. LANCET consists of two phases: (1) using influence functions to identify the training data that significantly impact undesirable model outputs, and (2) applying an Influence function-driven Bregman Optimization (IBO) technique to adjust the model's behavior based on these influence distributions. Our experiments demonstrate that LANCET effectively and efficiently correct inappropriate behaviors of LLMs. Furthermore, LANCET can outperform methods that rely on collecting human preferences, and it enhances the interpretability of learning human preferences within LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。