arXiv:2410.10014cs.CLcs.AI2024-10被引 47

自动识别并清除有害数据,提升大模型微调的安全性。

Safety-Aware Fine-Tuning of Large Language Models

  • 利用有害与良性样本的子空间特征设计评分函数
  • 在不同模型和污染率下,有害内容减少最高达27.8%
  • 适合关注模型安全性的实际应用开发者

大语言模型(LLMs)的微调已成为满足个性化需求的常用方法。然而,微调数据集来源多样,可能引入有害样本,手动筛选既耗时又主观。为此,我们提出一种新型安全感知微调(SAFT)框架,通过利用有害与良性样本的子空间信息设计评分函数,实现有害数据的自动检测与移除。实验表明,SAFT在不同大模型及不同污染率下均有效,可使有害性降低最多达27.8%。进一步分析揭示了该方法的机制,并验证其在真实场景中应对实际挑战的普适性。

原文摘要 · Abstract (English)

Fine-tuning Large Language Models (LLMs) has emerged as a common practice for tailoring models to individual needs and preferences. The choice of datasets for fine-tuning can be diverse, introducing safety concerns regarding the potential inclusion of harmful data samples. Manually filtering or avoiding such samples, however, can be labor-intensive and subjective. To address these difficulties, we propose a novel Safety-Aware Fine-Tuning (SAFT) framework designed to automatically detect and remove potentially harmful data, by leveraging a scoring function that exploits the subspace information of harmful and benign samples. Experimental results demonstrate the efficacy of SAFT across different LLMs and varying contamination rates, achieving reductions in harmfulness of up to 27.8%. Going beyond, we delve into the mechanism of our approach and validate its versatility in addressing practical challenges in real-world scenarios.

大模型安全微调数据清洗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。