arXiv:2502.11244cs.CLcs.AI2025-02EMNLP被引 13

通过微调特定语言的功能头,实现多语言模型的安全对齐。

Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment

  • 定位每种语言中生成有害内容的关键神经元,仅调整少量参数。
  • 在高、中、低资源语言中均显著降低违规率,且不影响模型性能。
  • 适合关注多语言安全对齐与伦理合规的研究者和开发者。

确保大语言模型在多语言环境下的安全一致性仍面临重大挑战。我们提出Soteria,一种轻量级但高效的方法,通过定位并最小化调整各语言中最可能导致有害内容生成的‘功能头’,仅修改少量参数即可大幅减少策略违规行为,且在低资源场景下仍保持模型整体性能。为严格评估该方法,我们还构建了XThreatBench,一个专门用于捕捉真实政策指南中的细粒度有害行为的多语言数据集。在主流开源LLM(如Llama、Qwen、Mistral)上的实验表明,Soteria在高、中、低资源语言中均能持续提升安全指标。这些发现揭示了一条通往全球可扩展、语言敏感且伦理对齐的大模型的可行路径。

原文摘要 · Abstract (English)

Ensuring consistent safety across multiple languages remains a significant challenge for large language models (LLMs). We introduce Soteria, a lightweight yet powerful strategy that locates and minimally adjusts the "functional heads" most responsible for harmful content generation in each language. By altering only a fraction of parameters, Soteria drastically reduces policy violations without sacrificing overall model performance, even in low-resource settings. To rigorously evaluate our approach, we also present XThreatBench, a specialized multilingual dataset capturing fine-grained harmful behaviors drawn from real policy guidelines. Experiments with leading open-source LLMs (e.g., Llama, Qwen, Mistral) show that Soteria consistently improves safety metrics across high-, mid-, and low-resource languages. These findings highlight a promising path toward scalable, linguistically attuned, and ethically aligned LLMs worldwide.

多语言安全参数微调伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。