arXiv:2502.02958cs.CL2025-02ICML被引 22

编辑大模型可能被恶意利用,威胁安全,亟需防护机制。

Position: Editing Large Language Models Poses Serious Safety Risks

  • 知识编辑技术易获取、低成本且隐蔽,易被滥用。
  • 可被用于篡改事实、制造虚假信息等恶意目的。
  • 适合关注AI安全、模型防护的研究者与政策制定者。

大型语言模型(LLMs)包含大量世界知识,但这些知识会随时间过时,催生了知识编辑方法(KEs),可在有限副作用下修改特定事实。本文指出,此类编辑存在严重安全隐患,长期被忽视。首先,KEs广泛可用、计算成本低、性能高且隐蔽,易被恶意使用者利用。其次,我们讨论了多种恶意应用场景,表明其可轻易适应于不同攻击目的。第三,当前AI生态存在漏洞,允许无验证地上传下载更新模型。第四,社会与制度层面缺乏认知,加剧了风险。我们呼吁学术界研究抗篡改模型及防御措施,并积极构建更安全的AI生态系统。

原文摘要 · Abstract (English)

Large Language Models (LLMs) contain large amounts of facts about the world. These facts can become outdated over time, which has led to the development of knowledge editing methods (KEs) that can change specific facts in LLMs with limited side effects. This position paper argues that editing LLMs poses serious safety risks that have been largely overlooked. First, we note the fact that KEs are widely available, computationally inexpensive, highly performant, and stealthy makes them an attractive tool for malicious actors. Second, we discuss malicious use cases of KEs, showing how KEs can be easily adapted for a variety of malicious purposes. Third, we highlight vulnerabilities in the AI ecosystem that allow unrestricted uploading and downloading of updated models without verification. Fourth, we argue that a lack of social and institutional awareness exacerbates this risk, and discuss the implications for different stakeholders. We call on the community to (i) research tamper-resistant models and countermeasures against malicious model editing, and (ii) actively engage in securing the AI ecosystem.

模型安全知识编辑恶意使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。