揭示大模型微调中能力与安全性的根本矛盾
Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models
- 构建理论框架分析微调时安全与能力的权衡机制
- 发现数据相似性、上下文重叠等影响权衡的关键因素
- 为提升安全性提供理论依据,适合模型安全研究者
在特定任务数据集上微调大型语言模型(LLMs)是其主要应用方式之一。然而,已有实证观察表明,这种提升能力的方法不可避免地损害安全性,即大模型微调中的安全-能力权衡现象。本文提出一个理论框架,用于理解两种主流安全感知微调策略中安全与能力的相互作用,揭示了数据相似性、上下文重叠以及对齐损失景观的影响。理论结果刻画了大模型微调中安全-能力权衡的根本极限,并通过数值实验得到验证。
原文摘要 · Abstract (English)
Fine-tuning Large Language Models (LLMs) on some task-specific datasets has been a primary use of LLMs. However, it has been empirically observed that this approach to enhancing capability inevitably compromises safety, a phenomenon also known as the safety-capability trade-off in LLM fine-tuning. This paper presents a theoretical framework for understanding the interplay between safety and capability in two primary safety-aware LLM fine-tuning strategies, providing new insights into the effects of data similarity, context overlap, and alignment loss landscape. Our theoretical results characterize the fundamental limits of the safety-capability trade-off in LLM fine-tuning, which are also validated by numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。