arXiv:2410.03772cs.CLcs.AI2024-10被引 1

精准修改模型有毒参数,提升大模型安全性

Precision Knowledge Editing: Enhancing Safety in Large Language Models

  • 通过神经元权重追踪与激活路径分析,定位并修正有毒内容区域
  • 在多个模型上将攻击成功率降至极低,同时保持原有性能
  • 适合关注大模型安全、可控生成的研究者与应用开发者

大型语言模型虽能力强大,但可能生成有害内容。本文提出精确知识编辑(PKE)技术,基于神经元权重追踪与激活路径追踪,比以往方法如DINM更精细地识别并修改模型中存在毒性的参数区域。实验表明,PKE显著降低多种模型(包括Llama2-7b和Llama-3-8b-instruct)的攻击成功率(ASR),同时维持整体性能。我们还对比了闭源模型gpt-4-0613与Claude 3 Sonnet,发现经本方法调整后的模型在安全性方面远超闭源模型。该研究为提升大模型在实际应用中的安全性和可靠性提供了有效手段。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities, but they also pose risks related to the generation of toxic or harmful content. This work introduces Precision Knowledge Editing (PKE), an advanced technique that builds upon existing knowledge editing methods to more effectively identify and modify toxic parameter regions within LLMs. By leveraging neuron weight tracking and activation pathway tracing, PKE achieves finer granularity in toxic content management compared to previous methods like Detoxifying Instance Neuron Modification (DINM). Our experiments demonstrate that PKE significantly reduces the attack success rate (ASR) across various models, including Llama2-7b and Llama-3-8b-instruct, while maintaining overall model performance. Additionally, we also compared the performance of some closed-source models (gpt-4-0613 and Claude 3 Sonnet) in our experiments, and found that models adjusted using our method far outperformed the closed-source models in terms of safety. This research contributes to the ongoing efforts to make LLMs safer and more reliable for real-world applications.

大模型安全知识编辑毒性控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。