用对比学习同时降低模型毒性并保持真实性和知识保留
Paying Alignment Tax with Contrastive Learning
- 通过构建正负样本对,用对比学习平衡去偏与忠实性
- 在多规模模型上实现毒性显著降低且知识保留率提升
- 首个能同时改善双指标的去偏方法,适合追求模型稳健性的研究者
当前去偏方法常导致模型能力下降,如事实准确性降低、知识丢失。我们在多个基准上系统评估发现,现有方法在小模型中存在根本性权衡,造成真实性减弱、知识流失或输出不可读。为此,我们提出一种对比学习框架,通过精心构造的正负样本对进行学习,并引入对比计算与动态损失缩放机制,以平衡去偏与忠实性。实验结果表明,该方法在多种模型规模下均显著提升毒性减少与忠实性保持效果。最重要的是,我们的框架是首个能持续同时优化这两项指标的方法,避免了现有方法的能力退化。结果表明,通过对比学习显式建模正负样本,可能是缓解语言模型去偏‘对齐税’的可行方向。
原文摘要 · Abstract (English)
Current debiasing approaches often result a degradation in model capabilities such as factual accuracy and knowledge retention. Through systematic evaluation across multiple benchmarks, we demonstrate that existing debiasing methods face fundamental trade-offs, particularly in smaller models, leading to reduced truthfulness, knowledge loss, or unintelligible outputs. To address these limitations, we propose a contrastive learning framework that learns through carefully constructed positive and negative examples. Our approach introduces contrast computation and dynamic loss scaling to balance bias mitigation with faithfulness preservation. Experimental results across multiple model scales demonstrate that our method achieves substantial improvements in both toxicity reduction and faithfulness preservation. Most importantly, we show that our framework is the first to consistently improve both metrics simultaneously, avoiding the capability degradation characteristic of existing approaches. These results suggest that explicit modeling of both positive and negative examples through contrastive learning could be a promising direction for reducing the alignment tax in language model debiasing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。