arXiv:2411.06824cs.AI2024-11被引 8

通过融合领域与对齐向量,提升专业大模型的安全性。

Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs

  • 用插值法合并领域和对齐向量,实现安全与专业性的平衡。
  • 在医疗和金融领域模型上,对齐性能显著提升,领域能力基本不变。
  • 适合需要安全可靠的行业专用大模型的研发人员参考。

目前,训练在特定技术领域表现优异的领域专家型大模型受到越来越多关注,但这类模型常伴随安全能力下降,可能生成有害内容。为此,我们提出一种高效且有效的基于合并的对齐方法——MergeAlign,通过插值领域向量与对齐向量,构建更安全的领域专用模型,同时保持其专业能力。我们在Llama3的医疗与金融领域变体上应用MergeAlign,实现显著的对齐性能提升,领域基准表现几乎无损。我们通过模型相似性度量与各模型贡献分析研究了合并的影响。希望本工作能开辟新研究方向,推动更高效安全的专家大模型开发。

原文摘要 · Abstract (English)

There is a growing interest in training domain-expert LLMs that excel in specific technical fields compared to their general-purpose instruction-tuned counterparts. However, these expert models often experience a loss in their safety abilities in the process, making them capable of generating harmful content. As a solution, we introduce an efficient and effective merging-based alignment method called \textsc{MergeAlign} that interpolates the domain and alignment vectors, creating safer domain-specific models while preserving their utility. We apply \textsc{MergeAlign} on Llama3 variants that are experts in medicine and finance, obtaining substantial alignment improvements with minimal to no degradation on domain-specific benchmarks. We study the impact of model merging through model similarity metrics and contributions of individual models being merged. We hope our findings open new research avenues and inspire more efficient development of safe expert LLMs.

大模型安全领域模型对齐方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。