arXiv:2606.28843cs.CLcs.AI2026-06

多语言微调可能让模型更易受攻击,且风险因语言而异。

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

论文配图:The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning
图 1 · 摘自论文原文
  • 用九种语言的普通数据微调大模型,观察安全影响。
  • 某些情况下,对抗性指令响应率飙升四倍,跨语言差异明显。
  • 英文外的语言微调后模型更易极端回应,需多语言评估。

微调大型语言模型是提升其在特定下游任务中能力的常用方法。然而,已有研究指出,这种能力提升会带来代价:即使使用非对抗性数据进行微调,也可能增加模型对不安全提示的响应倾向。本文首次在多语言环境下系统研究该现象,对 Llama-3.2、Qwen3 与 Gemma-3 模型使用九种语言的翻译数据进行微调。结果发现,安全表现高度依赖微调语言与评估语言,部分设置下对抗性指令合规率最高上升四倍。多语言安全漂移与通用能力指标解耦,且在不同语言和模型间呈现异质性。在非英语语言上微调通常导致较小的内部表征漂移,但反而使模型趋向过度顺从或完全拒绝。因此,仅以英语评估微调影响无法充分保障部署安全。为推动后续研究,我们发布了 Multilingual-Benign-Tune 数据集与 SORRY-Bench-Multilingual 评估套件。

原文摘要 · Abstract (English)

Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown that this increase in capability comes with a cost: it can increase a model's tendency to respond to unsafe adversarial prompts, even when fine-tuning with non-adversarial data. We present the first comprehensive empirical study of this phenomenon in multilingual settings by fine-tuning Llama-3.2, Qwen3, and Gemma-3 models using benign data translated across nine languages. We find that safety outcomes are highly sensitive to both the choice of fine-tuning language and the evaluation language, with adversarial compliance rates increasing four-fold in some settings. Multilingual safety drift is decoupled from general capability metrics, and occurs heterogeneously across languages and models. Fine-tuning in non-English languages often induces smaller internal representational drifts than English, but these shifts lead models to default to either exaggerated compliance or refusal. As such, assessing fine-tuning impacts solely in English provides inadequate assurance for deployment. To facilitate further research into these cross-lingual safety blind spots, we release the Multilingual-Benign-Tune dataset and the SORRY-Bench-Multilingual evaluation suite.

多语言安全风险微调模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。