arXiv:2608.25832cs.CLcs.AI2026-08

同一模型在不同语言下的表现差异显著,暴露了多语言能力的深层不平等。

Skill Issue: Are Skills Language-Invariant in LLMs?

论文配图:Skill Issue: Are Skills Language-Invariant in LLMs?
图 1 · 摘自论文原文
  • 用双语对战游戏隔离语言影响,量化模型跨语言技能差异。
  • 相同模型在八种语言中胜率差达20%以上,无效操作和策略偏差明显。
  • 调整推理语言可恢复部分性能,说明语言影响决策全过程。

大型语言模型在不同语言间知识获取不一致,其技能集是否存在跨语言差异?本文通过多语言自对弈框架,在固定规则与状态空间下,让同一模型以不同语言接口参与文本游戏,从而独立评估语言对行为的影响。我们在TextArena的多语言扩展上,评估了三个开源模型在八种语言、六类游戏(空间推理、不完全信息、资源分配、重复交互)中的表现。结果发现,同一模型在不同语言下的对战强度存在显著差异,胜率差距超过20%,无效操作频率与策略倾向也系统性变化。深入分析揭示了语言特异性错误:空间推理偏差、基于牌面条件的决策失误、最优动作选择失效。在某些场景中,仅改变中间推理语言即可恢复大部分性能损失,表明语言影响贯穿决策过程。该研究证明,技能不一致是构建真正多语言模型的重大障碍,理解此类差异有助于设计更公平的跨语言模型。

原文摘要 · Abstract (English)

Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance. We do this via multilingual self-play: two instances of the same model compete in a text-based game, each interacting through a different language interface. Since the model, opponent, rules, state space, and available actions remain fixed, this setting isolates the effect of language on the model's realized behavior. We build a multilingual extension to TextArena and evaluate three open-weight models across eight languages and six games covering spatial reasoning, imperfect information, resource allocation, and repeated interaction. We find that the same model can exhibit markedly different playing strength across languages, with systematic variation in win--loss margins, invalid actions, and strategic tendencies. Detailed analyses reveal language-specific failures in spatial reasoning, card-conditioned decisions, and optimal move selection. In some settings, changing only the intermediate reasoning language recovers much of the lost performance, suggesting that language can affect different stages of the decision process. These results show that skill discrepancies are a measurable major roadblock in the development of truly multilingual models. Better understanding these discrepancies can help us design models that perform more equitably across languages.

多语言模型技能不一致自对弈决策偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。