arXiv:2505.15508cs.CL2025-05被引 4

跨语言推理时缩放效果不一,提出方法提升低资源语言表现

Multilingual Test-Time Scaling via Initial Thought Transfer

  • 通过无监督前缀调优迁移高资源语言推理模板
  • 低资源语言初始推理一致性更低,且易中途转用英语
  • 特别改善小语种推理表现,适合多语言应用开发

测试时缩放已成为提升推理性能的常用推理阶段策略,但其有效性几乎仅在英语中被研究,其他语言中的表现尚不清楚。我们首次系统性地研究了多语言环境下测试时缩放的表现,评估了 DeepSeek-R1-Distill-LLama-8B 与 DeepSeek-R1-Distill-Qwen-7B 在高资源与低资源拉丁字母语言上的表现。结果表明,不同语言中测试时缩放带来的相对收益差异显著。此外,模型在执行单语提示时仍频繁在推理过程中切换至英语。我们进一步发现,低资源语言的初始推理思路与英语差异显著,且早期推理生成过程内部一致性较低。基于此,我们提出 MITT(Multilingual Initial Thought Transfer),一种无监督、轻量级的推理前缀调优方法,通过将高资源语言的推理前缀迁移至其他语言,以提升所有语言下的测试时缩放效果,缓解多语言推理性能不一致问题。MITT 显著提升了 DeepSeek-R1-Distill-Qwen-7B 在低资源语言上的推理表现。

原文摘要 · Abstract (English)

Test-time scaling has emerged as a widely adopted inference-time strategy for boosting reasoning performance. However, its effectiveness has been studied almost exclusively in English, leaving its behavior in other languages largely unexplored. We present the first systematic study of test-time scaling in multilingual settings, evaluating DeepSeek-R1-Distill-LLama-8B and DeepSeek-R1-Distill-Qwen-7B across both high- and low-resource Latin-script languages. Our findings reveal that the relative gains from test-time scaling vary significantly across languages. Additionally, models frequently switch to English mid-reasoning, even when operating under strictly monolingual prompts. We further show that low-resource languages not only produce initial reasoning thoughts that differ significantly from English but also have lower internal consistency across generations in their early reasoning. Building on our findings, we introduce MITT (Multilingual Initial Thought Transfer), an unsupervised and lightweight reasoning prefix-tuning approach that transfers high-resource reasoning prefixes to enhance test-time scaling across all languages, addressing inconsistencies in multilingual reasoning performance. MITT significantly boosts DeepSeek-R1-Distill-Qwen-7B's reasoning performance, especially for underrepresented languages.

多语言推理测试时缩放前缀调优低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。