arXiv:2504.01857cs.CLcs.AI2025-04被引 3

用多语言推理投票提升大模型逻辑一致性,显著改进小模型推理能力。

Cross-Lingual Consistency: A Novel Inference Framework for Advancing Reasoning in Large Language Models

  • 通过多语言推理路径集成与投票,缓解单语训练带来的偏见。
  • 在CMATH上使三款7B~9B模型准确率提升6%~9.5%,MGSM上提升4.1%~18.5%。
  • 适合关注多语言推理、模型鲁棒性与低资源场景的开发者和研究者。

思维链(CoT)已成为提升大语言模型(LLM)推理能力的关键机制,自一致性方法表现出显著潜力。然而,多语言训练语料中的固有语言偏见常导致语义漂移和逻辑不一致,尤其在参数量小于100亿的模型处理复杂推理任务时更为明显。为此,我们提出跨语言一致性(CLC)框架,一种创新的推理范式,通过多数投票整合多语言推理路径,以提升LLM的推理能力。在CMATH数据集上的实证评估显示,相较于传统自一致性方法,CLC在DeepSeek-Math-7B-Instruct、Qwen2.5-Math-7B-Instruct和Gemma2-9B-Instruct上分别实现9.5%、6.5%和6.0%的绝对准确率提升。将CLC的语言范围扩展至11种不同语言,带来双重优势:1)通过多语言集成投票中和多语言训练语料中的语言偏见;2)通过探索更广阔的多语言解空间,摆脱单语推理陷阱。实验表明,该方法相比单语自一致性基线,使Gemma2-9B-Instruct在MGSM数据集上实现4.1%至18.5%的准确率提升。

原文摘要 · Abstract (English)

Chain-of-thought (CoT) has emerged as a critical mechanism for enhancing reasoning capabilities in large language models (LLMs), with self-consistency demonstrating notable promise in boosting performance. However, inherent linguistic biases in multilingual training corpora frequently cause semantic drift and logical inconsistencies, especially in sub-10B parameter LLMs handling complex inference tasks. To overcome these constraints, we propose the Cross-Lingual Consistency (CLC) framework, an innovative inference paradigm that integrates multilingual reasoning paths through majority voting to elevate LLMs' reasoning capabilities. Empirical evaluations on the CMATH dataset reveal CLC's superiority over the conventional self-consistency method, delivering 9.5%, 6.5%, and 6.0% absolute accuracy gains for DeepSeek-Math-7B-Instruct, Qwen2.5-Math-7B-Instruct, and Gemma2-9B-Instruct respectively. Expanding CLC's linguistic scope to 11 diverse languages implies two synergistic benefits: 1) neutralizing linguistic biases in multilingual training corpora through multilingual ensemble voting, 2) escaping monolingual reasoning traps by exploring the broader multilingual solution space. This dual benefits empirically enables more globally optimal reasoning paths compared to monolingual self-consistency baselines, as evidenced by the 4.1%-18.5% accuracy gains using Gemma2-9B-Instruct on the MGSM dataset.

多语言推理思维链模型优化数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。