英语单语删去训练会导致多语言大模型产生语言混淆,影响评估效果。
Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
- 提出N-Mix评分量化多语言模型中的语言混淆现象
- 发现传统参考文本评估在高混淆时会误判为成功删去
- 建议采用直接分析生成内容的语义评估方法
现有研究指出,仅用英文数据尝试清除多语言大模型的知识是不足的。本文从评估视角切入,发现当多语言模型在全量平行语料上微调后进行删去训练时,会出现语言混淆现象——模型输出语言与输入提示语言不一致。该现象导致基于参考文本的标准评估指标失效。为此,我们提出三步方案:(1) 引入基于n-gram的N-Mix语言混杂评分,定量揭示语言混淆在多语言大模型中普遍存在且稳定;(2) 证明当N-Mix得分高时,参考基评估会产生假阴性;(3) 倡导引入直接分析生成内容语义的新评估方式,称之为语义基评估。本研究揭示了当前删去评估的盲区,推动更可靠的评估范式发展。
原文摘要 · Abstract (English)
There have been a couple of studies showing that attempting to erase multilingual knowledge using only English data is insufficient for multilingual LLMs. However, their analyses remain highly performance-oriented. In this paper, we switch the point of view to evaluation, and address an additional blind spot which reveals itself when the multilingual LLM is fully finetuned with parallel multilingual dataset before unlearning. Here, language confusion occurs whereby a model responds in language different from that of the input prompt. Language confusion is a problematic phenomenon in unlearning, causing the standard reference-based metrics to fail. We tackle this phenomenon in three steps: (1) introduce N-gram-based Language-Mix (N-Mix) score to quantitatively show the language confusion is pervasive and consistent in multilingual LLMs, (2) demonstrate that reference-based metrics result in false negatives when N-Mix score is high, and(3) suggest the need of new type of unlearning evaluation that can directly assess the content of the generated sentences. We call this type of metrics as semantic-based metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。