首次系统评估多语言思维链的性能、一致性与忠实性。
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
- 对比不同语言下模型思维链的质量与表现差异。
- 发现思维链在跨语言切换时效果显著变化。
- 适合关注多语言推理公平性与模型可信度的研究者。
大型推理模型越来越多地依赖分步思维链(CoT)来提升任务表现,尤其在英语等高资源语言中。尽管已有研究关注多语言场景下的最终答案准确率,但导致最终答案的中间推理过程——即思维链本身——仍缺乏深入探索。本文首次对多语言思维链进行系统评估,涵盖性能、一致性和忠实性三个维度。首先,通过显式指令或提示劫持让模型以目标语言思考,测量语言合规性、答案准确率与一致性,发现模型存在强烈语言偏好,跨语言表现差异显著。其次,通过交叉交换不同语言的思维链,评估其跨语言一致性,结果显示思维链质量与提示语言密切相关。最后,采用扰动技术(如截断和错误注入)探测各语言下思维链的忠实性,发现模型对思维链的依赖程度不一。代码与数据已开源,支持后续研究。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) increasingly rely on step-by-step Chain-of-Thought (CoT) reasoning to improve task performance, particularly in high-resource languages such as English. While recent work has examined final-answer accuracy in multilingual settings, the thinking traces themselves, i.e., the intermediate steps that lead to the final answer, remain underexplored. In this paper, we present the first comprehensive study of multilingual CoT reasoning, evaluating three key dimensions: performance, consistency, and faithfulness. We begin by measuring language compliance, answer accuracy, and answer consistency when LRMs are explicitly instructed or prompt-hacked to think in a target language, revealing strong language preferences and divergent performance across languages. Next, we assess crosslingual consistency of thinking traces by interchanging them between languages. We find that the quality and effectiveness of thinking traces vary substantially depending on the prompt language. Finally, we adapt perturbation-based techniques -- i.e., truncation and error injection -- to probe the faithfulness of thinking traces across languages, showing that models rely on traces to varying degrees. We release our code and data to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。