arXiv:2507.11544cs.CYcs.LG2025-07被引 4

开源大模型存在安全漏洞,该工具可评估其潜在危害差距

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models

  • 构建工具包量化评估开源模型在移除安全防护后的风险变化
  • 模型越大,安全缺口越明显,攻击能力提升显著
  • 适合安全研究者、模型开发者和政策制定者参考

开放权重的大语言模型在创新、个性化和隐私保护方面带来巨大优势,但其可修改性也引入系统性风险:恶意用户可轻易绕过现有安全机制,使有益模型变为有害工具。这形成了‘安全缺口’——即完整防护模型与被移除防护模型之间的危险能力差异。我们开源了安全缺口评估工具包,以两家族主流模型(Llama-3 和 Qwen-2.5)为对象,覆盖从 0.5B 到 405B 参数规模,采用多种去防护技术,评估其在生化与网络攻击能力、拒答率及生成质量上的表现。实验发现,随着模型规模增大,安全缺口持续扩大,去防护后危险能力显著增强。我们希望该工具包能成为通用开源模型的安全评估框架,并推动抗篡改防护机制的研发。欢迎社区共同贡献。

原文摘要 · Abstract (English)

Open-weight large language models (LLMs) unlock huge benefits in innovation, personalization, privacy, and democratization. However, their core advantage - modifiability - opens the door to systemic risks: bad actors can trivially subvert current safeguards, turning beneficial models into tools for harm. This leads to a 'safety gap': the difference in dangerous capabilities between a model with intact safeguards and one that has been stripped of those safeguards. We open-source a toolkit to estimate the safety gap for state-of-the-art open-weight models. As a case study, we evaluate biochemical and cyber capabilities, refusal rates, and generation quality of models from two families (Llama-3 and Qwen-2.5) across a range of parameter scales (0.5B to 405B) using different safeguard removal techniques. Our experiments reveal that the safety gap widens as model scale increases and effective dangerous capabilities grow substantially when safeguards are removed. We hope that the Safety Gap Toolkit (https://github.com/AlignmentResearch/safety-gap) will serve as an evaluation framework for common open-source models and as a motivation for developing and testing tamper-resistant safeguards. We welcome contributions to the toolkit from the community.

模型安全开源模型风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。