arXiv:2512.22712cs.CL2025-12中稿 · EMNLP被引 2

发现大模型跨语言推理常自相矛盾,非拉丁语系问题更严重。

Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages

  • 构建人类验证框架,检测模型推理是否支持结论。
  • 65k条跨语言推理中,非拉丁语系推理与结论错位率超拉丁语系两倍。
  • 揭示推理错误主因是证据不足和逻辑断裂,适合评估多语言模型的团队参考。

大型语言模型通过思维链提示展现出强大的推理能力,但其推理质量在不同语言间的迁移性仍缺乏深入研究。本文提出一种人类验证框架,用于评估模型生成的推理过程是否逻辑上支持其最终结论。通过对全球MMLU数据集中6种语言、6个前沿模型生成的6.5万条推理轨迹进行分析,我们发现一个关键盲点:尽管模型任务准确率高,其推理过程却常常无法支持最终结论。使用非拉丁字母书写语言的推理轨迹,其推理与结论之间的不一致程度至少是拉丁字母语言的两倍。通过人工标注,我们构建了一个错误分类体系,发现错误主要源于证据性错误(无依据断言、模糊事实)以及不合逻辑的推理步骤。研究结果表明,当前多语言评估方法未能全面反映模型的真实推理能力,亟需引入以推理为核心的评估框架。

原文摘要 · Abstract (English)

Large language models demonstrate strong reasoning capabilities through chain-of-thought prompting, but whether this reasoning quality transfers across languages remains underexplored. We introduce a human-validated framework to evaluate whether model-generated reasoning traces logically support their conclusions across languages. Analyzing 65k reasoning traces from GlobalMMLU questions across 6 languages and 6 frontier models, we uncover a critical blind spot: while models achieve high task accuracy, their reasoning can fail to support their conclusions. Reasoning traces in non-Latin scripts show at least twice as much misalignment between their reasoning and conclusions than those in Latin scripts. We develop an error taxonomy through human annotation to characterize these failures, finding they stem primarily from evidential errors (unsupported claims, ambiguous facts) followed by illogical reasoning steps. Our findings demonstrate that current multilingual evaluation practices provide an incomplete picture of model reasoning capabilities and highlight the need for reasoning-aware evaluation frameworks.

多语言推理评估模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。