arXiv:2506.01840cs.CL2025-06ACL被引 11

用最小差异对评估大模型代码转换能力,发现越大模型越像双语者。

Minimal Pair-Based Evaluation of Code-Switching

  • 构造自然与微调的代码转换句子对,对比偏好差异。
  • 大模型越大,越倾向选择自然的代码转换句,尤其在封闭词上差异显著。
  • 适用于研究双语行为与大模型语言理解的学者。

目前缺乏一种评估方法,可衡量大语言模型(LLMs)在代码转换(CS)中是否像双语者一样使用。现有方法存在语言覆盖范围窄、无法涵盖多样化的代码转换现象或难以扩展的问题。本文提出基于最小差异对的干预方法:每对包含一个自然发生的代码转换句子及其微小修改的变体。我们为11组语言对分别收集了最多1000个这样的配对。人类实验显示,对每组语言对,双语者均一致偏好自然的代码转换句。而当前大模型的实验表明,模型越大,越倾向于为自然句子分配更高的概率。根据理论预期,概率差异最大的情况出现在修改部分为封闭类词的语言对中。

原文摘要 · Abstract (English)

There is a lack of an evaluation methodology that estimates the extent to which large language models (LLMs) use code-switching (CS) in the same way as bilinguals. Existing methods do not have wide language coverage, fail to account for the diverse range of CS phenomena, or do not scale. We propose an intervention based on minimal pairs of CS. Each minimal pair contains one naturally occurring CS sentence and one minimally manipulated variant. We collect up to 1,000 such pairs each for 11 language pairs. Our human experiments show that, for every language pair, bilinguals consistently prefer the naturally occurring CS sentence. Meanwhile our experiments with current LLMs show that the larger the model, the more consistently it assigns higher probability to the naturally occurring CS sentence than to the variant. In accordance with theoretical claims, the largest probability differences arise in those pairs where the manipulated material consisted of closed-class words.

代码转换大模型双语研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。