用KTO方法让低资源语言模型更安全,效果提升99%。
Safe at the Margins: A General Approach to Safety Alignment in Low-Resource English Languages -- A Singlish Case Study
- 结合SFT与KTO优化模型毒性
- 毒性降低99%,且样本效率更高
- 适合多语言安全对齐研究者
确保大语言模型在多样化语言环境中的安全性仍具挑战,尤其针对低资源语言。现有安全对齐方法以英语为中心,效果受限。我们系统比较了监督微调(SFT)、直接偏好优化(DPO)和卡尼曼-特韦尔斯基优化(KTO)对SEA-Lion-v2.1-Instruct(Llama 3-8B变体)的毒性抑制能力。结果表明,SFT+KTO在样本效率上优于DPO,实现更优的安全对齐。此外,我们提出KTO-S,通过改进KL散度正则化增强稳定性。该方法将Singlish毒性降低99%,并可泛化至TOXIGEN数据集,在标准LLM基准测试中保持高性能,为多语言环境下安全AI部署提供可扩展框架。
原文摘要 · Abstract (English)
Ensuring the safety of Large Language Models (LLMs) in diverse linguistic settings remains challenging, particularly for low-resource languages. Existing safety alignment methods are English-centric, limiting their effectiveness. We systematically compare Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Kahneman-Tversky Optimization (KTO) for aligning SEA-Lion-v2.1-Instruct, a Llama 3-8B variant, to reduce toxicity in Singlish. Our results show that SFT+KTO achieves superior safety alignment with higher sample efficiency than DPO. Additionally, we introduce KTO-S, which enhances stability via improved KL divergence regularization. Our approach reduces Singlish toxicity by 99\%, generalizes to TOXIGEN, and maintains strong performance on standard LLM benchmarks, providing a scalable framework for safer AI deployment in multilingual contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。