arXiv:2505.22298cs.CL2025-05ACL被引 8

提出动态检测毒性激活的编辑方法,精准去毒不伤模型通用能力。

Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing

  • 通过前向传播中动态识别毒性模式,智能选择处理路径。
  • 在多个大模型上实现更强去毒效果,同时减少对正常指令的误拒。
  • 改进评测基准,更真实反映模型安全与能力平衡问题。

大型语言模型虽具备强大语言能力,但仍易受恶意提示和越狱攻击影响。现有知识编辑去毒方法存在两大挑战:一是依赖特定实体定位,对无显式实体的对抗输入无效;二是过度编辑问题,导致模型拒绝合法查询,损害整体性能。本文提出ToxEdit,一种毒性感知的知识编辑方法,可在前向传播中动态检测毒性激活模式,并通过自适应层间路径路由实现有效抑制。该设计确保精准去毒的同时保留模型通用能力。为更准确评估过度编辑问题,我们还改进了SafeEdit基准,加入指令遵循类评估任务。在多个大模型上的实验表明,ToxEdit在去毒性能和保护通用能力方面均优于现有最先进方法。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit impressive language capabilities but remain vulnerable to malicious prompts and jailbreaking attacks. Existing knowledge editing methods for LLM detoxification face two major challenges. First, they often rely on entity-specific localization, making them ineffective against adversarial inputs without explicit entities. Second, these methods suffer from over-editing, where detoxified models reject legitimate queries, compromising overall performance. In this paper, we propose ToxEdit, a toxicity-aware knowledge editing approach that dynamically detects toxic activation patterns during forward propagation. It then routes computations through adaptive inter-layer pathways to mitigate toxicity effectively. This design ensures precise toxicity mitigation while preserving LLMs' general capabilities. To more accurately assess over-editing, we also enhance the SafeEdit benchmark by incorporating instruction-following evaluation tasks. Experimental results on multiple LLMs demonstrate that our ToxEdit outperforms previous state-of-the-art methods in both detoxification performance and safeguarding general capabilities of LLMs.

大模型安全知识编辑去毒鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。