arXiv:2511.08484cs.AI2025-11被引 1

用微小补丁快速修复大模型安全漏洞,不改主模型也能提升安全性。

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models

  • 在现有模型前加可学习的微型前缀作为补丁
  • 仅增加0.003%参数,就能达到新一代安全模型效果
  • 适合需要快速迭代安全策略的厂商和开发者

我们提出将大语言模型(LLM)的安全更新类比为软件补丁,一种轻量且模块化的安全漏洞修复方法。尽管厂商会发布更安全的新版本模型,但重大更新成本高、频率低,难以满足客户个性化需求,导致已发布模型仍存在已知安全缺陷。与全模型微调或大型版本升级不同,我们的方法通过在现有模型前添加一个紧凑、可学习的前缀实现快速修复。该‘补丁’仅引入0.003%额外参数,却能可靠地引导模型行为向更安全的参考模型靠拢。在毒性缓解、偏见减少和有害内容拒绝三个关键领域,政策补丁均实现了与下一代对齐模型相当的安全性提升,同时保持语言流畅性。结果表明,大模型可像软件一样被‘打补丁’,为厂商和从业者提供了在重大发布之间分发可扩展、高效且可组合的安全更新的实用机制。

原文摘要 · Abstract (English)

We propose patching for large language models (LLMs) like software versions, a lightweight and modular approach for addressing safety vulnerabilities. While vendors release improved LLM versions, major releases are costly, infrequent, and difficult to tailor to customer needs, leaving released models with known safety gaps. Unlike full-model fine-tuning or major version updates, our method enables rapid remediation by prepending a compact, learnable prefix to an existing model. This "patch" introduces only 0.003% additional parameters, yet reliably steers model behavior toward that of a safer reference model. Across three critical domains (toxicity mitigation, bias reduction, and harmfulness refusal) policy patches achieve safety improvements comparable to next-generation safety-aligned models while preserving fluency. Our results demonstrate that LLMs can be "patched" much like software, offering vendors and practitioners a practical mechanism for distributing scalable, efficient, and composable safety updates between major model releases.

大模型安全模型补丁轻量更新行为对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。