让低资源东南亚语言模型稳定进行多语言推理,避免英文干扰。
Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages

- 用翻译代理循环将高资源语言推理路径映射到低资源语言空间。
- 在数学推理上提升最高3.2倍,语言偏见减少6.4%。
- 适合做东南亚小语种推理任务的研究者和应用开发者。
大型语言模型在推理能力上取得显著进展,但在低资源原生语境下,常因跨语言坍缩而回退至英语进行复杂逻辑推理,导致政策优化的冷启动瓶颈。标准微调又面临灾难性遗忘风险,源于跨语言表征漂移。为此,我们提出基于引导序列的交叉蒸馏(OSCD)后训练算法:通过集成翻译代理循环,在生成式训练推演中将高资源语言推理轨迹投影至低资源语言词汇子空间,确保动态生成参考样本的稳定高效翻译以用于微调。同时联合嵌入对齐参考语言与目标语言推理轨迹的语义,弥合跨语言表征差距。在AIME25和HMMT25基准上的综合评估显示,OSCD使东南亚原生语言数学推理性能最高提升3.2倍,其中联合嵌入语义对齐组件相较仅翻译基线提升最多6.4%的语言去偏效果。
原文摘要 · Abstract (English)
Large Language Models have achieved substantial progress in reasoning capabilities. Yet in low-resource native settings, many suffer from cross-lingual collapse, reverting to English during intermediate steps that require complex logical reasoning. This presents a cold-start bottleneck for policy optimization, whereas standard fine-tuning risks catastrophic forgetting due to cross-lingual representation drift. To address these challenges, we introduce the Onramp-Sequence Cross-Distillation (OSCD), a post-training algorithm that projects high-resource reasoning trajectories into low-resource vocabulary subspaces during generative training rollouts via an integrated translator agentic loop, ensuring the stable and efficient translation of dynamically generated reference samples for fine-tuning. This is coupled with joint-embedding semantic alignment of both reference and target-language reasoning traces, thereby bridging the pairwise cross-lingual representational gaps. Comprehensive evaluations using the AIME25 and HMMT25 benchmarks demonstrate that OSCD yields up to 3.2 times overall improvements in native Southeast Asian languages for mathematical reasoning, of which the joint-embedding semantic alignment component contributes up to 6.4% improvements in linguistic debiasing over translation-only baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。