检测知识编辑方法在去毒中的真实效果,发现多数看似有效实则虚假。
On the Robustness of Knowledge Editing for Detoxification
- 从优化、组合和跨语言三方面评估去毒可靠性
- 发现编辑后毒性降低可能是生成异常而非真正去毒
- 仅特定模型、少量目标和部分语言下方法才可靠
基于知识编辑的去毒方法被视作缓解大模型有害行为的有前景方案。然而现有评估多依赖自动毒性分类器,隐含假设毒性分数下降即代表行为真正抑制。本文提出面向鲁棒性的评估框架,从优化鲁棒性、组合鲁棒性和跨语言鲁棒性三个维度检验该方法可靠性。我们识别出一种常见失效模式——伪去毒:表面毒性降低源于退化生成行为,而非对不当内容的真实抑制。进一步发现,当多个有害行为同时编辑时,去毒效果显著下降;且单语和跨语言去毒仅在特定模型-方法组合下有效。总体表明,基于知识编辑的去毒仅在特定模型、有限目标数量及部分语言中具备鲁棒性。
原文摘要 · Abstract (English)
Knowledge-Editing-based (KE-based) detoxification has emerged as a promising approach for mitigating harmful behaviours in Large Language Models. Existing evaluations, however, largely rely on automatic toxicity classifiers, implicitly assuming that reduced toxicity scores reflect genuine behavioural suppression. In this work, we propose a robustness-oriented evaluation framework for KE-based detoxification that examines its reliability beyond standard classifier-based metrics along three dimensions: optimisation robustness, compositional robustness, and cross-lingual robustness. We identify pseudo-detoxification as a common failure mode, where apparent toxicity reductions arise from degenerate generation behaviours rather than meaningful suppression of unsafe content. We further show that detoxification effectiveness degrades when multiple unsafe behaviours are edited jointly, and that both monolingual and cross-lingual detoxification remain effective only under specific model-method combinations. Overall, our results indicate that KE-based detoxification is robust only for certain models, limited numbers of detoxification objectives, and a subset of languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。