用分层替换技术让低资源语言模型更安全,无需额外训练。
Layer-wise Swapping for Generalizable Multilingual Safety
- 通过分层替换将英语安全专家的知识迁移到低资源语言模型。
- 在多语言安全测试中,生成内容更合规且危害性更低。
- 适合需要低成本提升多语言安全性的AI研发团队使用。
尽管大型语言模型进展迅速,低资源语言的安全风险仍是关键挑战。现有安全数据集以英语为主,限制了多语言安全对齐的进展。因此,针对特定语言微调的低资源专家模型,其不安全性远高于高资源模型。本文提出一种安全感知的分层替换方法,无需额外训练即可将英语安全专家的对齐能力迁移至低资源语言模型。为增强迁移效果,方法根据模块专业化程度自适应选择或融合。实验表明,该方法在通用基准如MMMLU、BELEBELE和MGSM上表现接近原语言专家,同时在MultiJail安全基准上生成更合规、危害性更低的响应。
原文摘要 · Abstract (English)
Despite the rapid advancements of Large Language Models (LLMs), safety risks remain a critical challenge for low-resource languages. Existing safety datasets are predominantly English centric, limiting progress in multilingual safety alignment. As a result, low resource expert models, finetuned on their respective instruction datasets, tend to exhibit higher unsafety rates compared to their high resource counterparts. In this work, we propose a safety aware layer swapping method that transfers safety alignment from an English safety expert to low resource language experts without additional training. To further enhance transfer ability, our method adaptively selects or blends modules based on their degree of specialization. Our approach preserves performance on general language understanding tasks while enhancing safety in the target languages. Experimental results show that the proposed method achieves comparable performance to the language expert on general benchmarks such as MMMLU, BELEBELE, and MGSM, while producing more aligned and less harmful responses on the MultiJail safety benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。