arXiv:2608.08447cs.CLcs.AI2026-08

研究大模型推理时语言一致性随难度变化的规律,发现高难度下语言会突然切换。

Hidden Language Consistency Phenomena in Reasoning LLMs

论文配图:Hidden Language Consistency Phenomena in Reasoning LLMs
图 1 · 摘自论文原文
  • 分析八种语言在四个难度等级下的推理语言一致性表现
  • 高难度时语言一致性骤降,尤其非拉丁字母语言更明显
  • 适合关注多语言模型真实能力评估的研究者

多语言推理模型通常只评估答案正确性,却忽略推理与回应过程中是否保持原始语言。这种忽略掩盖了任务变难时出现的重要多语言行为。本文基于PolyMath基准,在八种语言和四个难度层级上研究了任务难度、准确率、思维语言一致性(TC)与回答语言一致性(AC)。发现四类语言一致性行为:输出语言一致性可能保持一致、持续错位、缓慢退化或突然崩溃;识别出语言一致性崩溃现象,即难度上升会导致输出语言一致性突然下降,尤其在低频和非拉丁文字语言中更显著;由于此崩溃效应,模型转用内部主导语言后,准确率甚至可能提升;量化方法可独立影响语言一致性,如GPTQ和AWQ在ε=1.0容忍投票下常优于AutoRound。结果表明,多语言能力不能仅靠准确率衡量,评估需综合考虑准确率、语言一致性与任务难度。

原文摘要 · Abstract (English)

Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding. This omission conceals important multilingual behaviors that emerge as tasks become harder. In this paper, we study task difficulty, task accuracy, thinking-language consistency (TC), and answer-language consistency (AC) across reasoning models using PolyMath benchmark in eight languages and four difficulty levels. We uncover four findings: (1) language consistency exhibits four difficulty-dependent behaviors: output-language consistency remains aligned with input, remains misaligned, degrades gradually, or collapses abruptly. (2) We identify the language consistency breakdown effect, where increasing difficulty can cause a sudden drop in output-language consistency, especially in less strongly represented and non-Latin-script languages. (3) Due to this breakdown effect, accuracy can be preserved or even improved at a harder difficulty level as the model shifts to its internal dominant language. (4) Quantization can improve or degrade output-language consistency independently of its effect on accuracy, with GPTQ and AWQ often outperforming AutoRound under tolerance-based voting with ε = 1.0. These results show that multilingual capability cannot be characterized by accuracy alone; reliable evaluation should jointly consider task accuracy, language consistency, and task difficulty for multilingual benchmarks.

多语言模型语言一致性推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。