用英语安全信号指导多语言模型拒绝有害请求,提升跨语言安全性。
BabelSteering: Multilingual Safety Alignment via English Steering Vectors

- 通过英语安全方向向量,在推理时轻量干预激活值以增强多语言拒答能力。
- 多语言测试中,有害请求拒答率平均提升11个百分点,任务性能几乎不变。
- 适合关注多语言AI安全、需低成本部署的开发者和研究者。
大型语言模型在高风险场景中全球部署,但多数安全研究集中于英语。本研究探索是否可将英语安全信号迁移到其他语言。提出BabelSteering,一种基于激活值的轻量级推理时干预方法,利用英语安全监督提取的拒绝方向,实现跨语言泛化。评估涵盖八种语言,综合衡量有害请求拒答率、过度拒答及通用任务效用。结果表明,该方法显著提升多语言有害请求拒答率,平均提高11个百分点(如孟加拉语达17个百分点),对Global MMLU任务性能无影响,但伪有害提示拒答率平均上升13个百分点。同时构建多语言翻译与评估流水线,助力未来跨语言安全研究。研究显示,激活转向是扩展英语安全信号至多语言的实用且低成本方案。
原文摘要 · Abstract (English)
Large language models (LLMs) are deployed globally in high-stakes settings, yet most safety research and alignment efforts remain concentrated on English. Thus, users interacting with LLMs in other languages may encounter weaker safeguards despite relying on the same systems for similarly sensitive tasks. In this work, we investigate whether safety signals learned from a high-resource language, like English, can improve multilingual safety. We propose BabelSteering, an activation steering method that acts as a lightweight inference- time intervention, using refusal directions derived from English safety supervision to generalize across languages. Our evaluation includes eight languages and jointly measures refusal of harmful requests, over-refusal, and general task utility. The results show that BabelSteering increases the refusal of harmful requests across languages, with only a marginal to no reduction in task utility but with some increase in refusal of pseudo-harmful prompts. For example, for Gemma 7B, we see an average increase in the refusal of harmful prompts across languages of 11 percentage points (pp), with individual languages like Bengali seeing an increase of 17 pp, with no loss of utility on Global MMLU, while pseudo-harmful refusals increase by 13 pp on average. We also introduce a multilingual translation-and-evaluation pipeline to facilitate future work on cross-lingual safety interventions. Overall, our findings suggest that activation steering may provide a practical, low- cost mechanism for extending English-derived safety signals to other languages. Warning: this paper contains examples with unsafe content
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。