用英语推理模型跨语言解题,高资源语言效果更好。
Crosslingual Reasoning through Test-Time Scaling
- 通过扩大推理计算量,提升英语模型跨语言数学推理能力。
- 模型在高资源语言中推理更准确高效,低资源语言仍不足。
- 发现模型会自动用引述+思考模式处理非英语输入,可控制语言输出。
大型语言模型的推理能力主要针对英语研究,尽管预训练模型具备多语言能力。本文探究以英语长链思维(CoT)微调的推理模型在跨语言场景下的泛化能力。首先发现,增加英语中心推理模型的推理计算量,可在多种语言(包括低资源语言)上显著提升数学推理表现,甚至超越规模两倍的模型。其次,尽管英文推理提示以英语为主,但模型会采用‘引用+思考’模式处理非英语输入,实现跨语言推理。第三,提出有效策略控制长思维链的语言,结果显示模型在高资源语言中推理更优且更高效。最后,发现即使在英语内部,从理工科到文化常识的跨领域推理仍表现不佳。研究揭示了英语推理测试时扩展的潜力、机制与局限,建议实践者让英语模型在高资源语言中推理,同时需进一步改进低资源语言和跨领域推理能力。
原文摘要 · Abstract (English)
Reasoning capabilities of large language models are primarily studied for English, even when pretrained models are multilingual. In this work, we investigate to what extent English reasoning finetuning with long chain-of-thoughts (CoTs) can generalize across languages. First, we find that scaling up inference compute for English-centric reasoning language models (RLMs) improves multilingual mathematical reasoning across many languages including low-resource languages, to an extent where they outperform models twice their size. Second, we reveal that while English-centric RLM's CoTs are naturally predominantly English, they consistently follow a quote-and-think pattern to reason about quoted non-English inputs. Third, we discover an effective strategy to control the language of long CoT reasoning, and we observe that models reason better and more efficiently in high-resource languages. Finally, we observe poor out-of-domain reasoning generalization, in particular from STEM to cultural commonsense knowledge, even for English. Overall, we demonstrate the potentials, study the mechanisms and outline the limitations of crosslingual generalization of English reasoning test-time scaling. We conclude that practitioners should let English-centric RLMs reason in high-resource languages, while further work is needed to improve reasoning in low-resource languages and out-of-domain contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。