arXiv:2511.15886cs.CL2025-11被引 1

分析多语言大模型推理链的可信度,发现最终步骤被过度强调。

What Really Counts? Examining Step and Token Level Attribution in Multilingual CoT Reasoning

  • 用两种方法分析推理步骤和词语层面的贡献度
  • 错误生成中最终步骤的得分异常偏高
  • 高资源拉丁语系语言受益更明显,跨语言表现不稳

本研究考察多语言大模型在链式思维(CoT)推理中的归因模式。尽管先前工作表明CoT提示能提升任务性能,但其生成推理链的可信度与可解释性仍存疑。为评估跨语言特性,我们采用两种互补的归因方法——ContextCite用于步骤级归因,Inseq用于词级归因——对Qwen2.5 1.5B-Instruct模型在MGSM基准上的表现进行分析。实验结果表明:(1)归因分数过度聚焦于最终推理步骤,尤其在错误生成中更为显著;(2)结构化CoT提示主要显著提升高资源拉丁字母语言的准确率;(3)通过否定句和干扰句进行可控扰动后,模型准确率与归因一致性均下降。这些发现揭示了CoT提示在多语言鲁棒性和解释透明性方面的局限性。

原文摘要 · Abstract (English)

This study investigates the attribution patterns underlying Chain-of-Thought (CoT) reasoning in multilingual LLMs. While prior works demonstrate the role of CoT prompting in improving task performance, there are concerns regarding the faithfulness and interpretability of the generated reasoning chains. To assess these properties across languages, we applied two complementary attribution methods--ContextCite for step-level attribution and Inseq for token-level attribution--to the Qwen2.5 1.5B-Instruct model using the MGSM benchmark. Our experimental results highlight key findings such as: (1) attribution scores excessively emphasize the final reasoning step, particularly in incorrect generations; (2) structured CoT prompting significantly improves accuracy primarily for high-resource Latin-script languages; and (3) controlled perturbations via negation and distractor sentences reduce model accuracy and attribution coherence. These findings highlight the limitations of CoT prompting, particularly in terms of multilingual robustness and interpretive transparency.

链式思维多语言可解释性归因分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。