用数据修正法降低大模型生成内容的有害性与越狱风险。
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content
- 通过后生成修正机制,动态调整输出内容安全性。
- 多模型测试显示毒性与越狱分数平均下降15%以上。
- 无需微调模型,适合实际部署中的安全防护需求。
大型语言模型(LLM)虽取得显著进展,但潜在偏见和有害内容仍引发担忧。为此,我们提出一种实用方案,确保LLM的安全与伦理使用。新方法聚焦于后生成修正机制——BART-Corrective Model,通过调整生成内容提升安全性。相比仅依赖模型微调或提示工程,该方法提供了一种稳健的数据中心替代方案。在多个有毒文本数据集上的实验表明,集成该方法后,均值毒性与越狱得分显著降低:GPT-4分别减少15%和21%,PaLM2减少28%和5%,Mistral-7B约减少26%和23%,Gemma-2b-it减少11.1%和19%。结果证明该方法能有效提升LLM的安全性与可靠性,更适用于真实场景应用。
原文摘要 · Abstract (English)
Large Language Models (LLM) have made remarkable progress, but concerns about potential biases and harmful content persist. To address these apprehensions, we introduce a practical solution for ensuring LLM's safe and ethical use. Our novel approach focuses on a post-generation correction mechanism, the BART-Corrective Model, which adjusts generated content to ensure safety and security. Unlike relying solely on model fine-tuning or prompt engineering, our method provides a robust data-centric alternative for mitigating harmful content. We demonstrate the effectiveness of our approach through experiments on multiple toxic datasets, which show a significant reduction in mean toxicity and jail-breaking scores after integration. Specifically, our results show a reduction of 15% and 21% in mean toxicity and jail-breaking scores with GPT-4, a substantial reduction of 28% and 5% with PaLM2, a reduction of approximately 26% and 23% with Mistral-7B, and a reduction of 11.1% and 19% with Gemma-2b-it. These results demonstrate the potential of our approach to improve the safety and security of LLM, making them more suitable for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。