针对细粒度毒性,动态选择专用探针向量实现精准去毒。
DAPI: Domain Adaptive Toxicity Probe Vector Intervention for Fine-Grained Detoxification
- 为不同毒性类别训练专用探针向量,生成时按上下文动态选择。
- 在评估数据集上毒性降低78.52%,流畅度仅下降0.052%。
- 适合需要精细控制文本安全性的生成模型应用。
现有研究多采用单一毒性探针向量进行去毒,但毒性具有细粒度特征,单一向量难以覆盖所有类型。为此,我们提出类别特异性毒性探针向量方法:首先为不同毒性类别训练多个探针向量;生成时根据当前上下文动态选择最相关向量;最后动态缩放并从模型中减去该向量。实验表明,该方法成功去除了单向量方法无法处理的毒性类别,在评估数据集上实现最高78.52%的毒性降低,同时流畅度仅比未引导模型下降0.052%,保持了生成质量。
原文摘要 · Abstract (English)
There have been attempts to utilize linear probe for detoxification, with existing studies relying on a single toxicity probe vector to reduce toxicity. However, toxicity can be fine-grained into various subcategories, making it difficult to remove certain types of toxicity by using a single toxicity probe vector. To address this limitation, we propose a category-specific toxicity probe vector approach. First, we train multiple toxicity probe vectors for different toxicity categories. During generation, we dynamically select the most relevant toxicity probe vector based on the current context. Finally, the selected vector is dynamically scaled and subtracted from model. Our method successfully mitigated toxicity from categories that the single probe vector approach failed to detoxify. Experiments demonstrate that our approach achieves up to a 78.52% reduction in toxicity on the evaluation dataset, while fluency remains nearly unchanged, with only a 0.052% drop compared to the unsteered model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。