arXiv:2510.13183cs.CL2025-10EMNLP被引 4

无需微调的轻量级去毒方法,提升大模型输出安全性和流畅性。

DSCD: Large Language Model Detoxification with Self-Constrained Decoding

  • 利用自约束解码动态调整生成过程中的安全、幻觉与有毒层分布。
  • 在多个开源模型上实现当前最优的去毒效果与生成流畅度。
  • 无需参数修改,可无缝集成现有去毒工具,适合实际部署场景。

大语言模型的去毒仍是重大研究挑战。现有解码去毒方法均依赖外部约束,需额外资源开销且影响生成流畅性。本文提出无需参数微调的自约束解码(DSCD)方法,通过增强生成过程中安全层的词元分布,同时削弱幻觉和有毒层的分布,有效降低毒性并提升输出安全性。DSCD具有轻量化、高兼容性和即插即用特性,可与现有去毒方法无缝结合以进一步提升性能。在代表性开源大模型和公开数据集上的大量实验验证了其有效性,展现了当前最优的去毒性能与生成流畅性,且效率优于已有方法。结果表明DSCD是更实用、可扩展的大模型安全部署方案。

原文摘要 · Abstract (English)

Detoxification in large language models (LLMs) remains a significant research challenge. Existing decoding detoxification methods are all based on external constraints, which require additional resource overhead and lose generation fluency. This work proposes Detoxification with Self-Constrained Decoding (DSCD), a novel method for LLM detoxification without parameter fine-tuning. DSCD strengthens the inner next-token distribution of the safety layer while weakening that of hallucination and toxic layers during output generation. This effectively diminishes toxicity and enhances output safety. DSCD offers lightweight, high compatibility, and plug-and-play capabilities, readily integrating with existing detoxification methods for further performance improvement. Extensive experiments on representative open-source LLMs and public datasets validate DSCD's effectiveness, demonstrating state-of-the-art (SOTA) performance in both detoxification and generation fluency, with superior efficiency compared to existing methods. These results highlight DSCD's potential as a practical and scalable solution for safer LLM deployments.

大模型安全去毒解码优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。