arXiv:2604.21593cs.CL2026-04被引 2

让语言成为推理优化的隐变量,提升模型跨任务表现

Language as a Latent Variable for Reasoning Optimization

论文配图:Language as a Latent Variable for Reasoning Optimization
图 1 · 摘自论文原文
  • 将语言视为隐变量,通过多语言约束与自由生成对比优化推理路径
  • 在仅18.1K数学题上训练,英文推理准确率提升6.72%,多语言提升6.89%
  • 仅用数学数据训练却超越基线模型在英语常识推理上的表现

随着大语言模型减少对英语的依赖,一个令人意外的现象出现:非英语回答在推理任务中有时优于英语。我们假设语言作为潜在变量,会结构性地调节模型内部推理路径,而非仅作为输出媒介。为此,我们开展多语言思考实验,在语言受限与无限制条件下让模型解决相同问题。结果显示,非英语回答常更准确,最佳表现多出现在语言无约束时,表明多语言性拓宽了模型的潜在推理空间。基于此,我们提出polyGRPO(多语言组相对策略优化)框架,将语言差异作为隐式探索信号,在线生成多语言偏好数据,同时优化答案准确率与推理结构。在仅18.1K多语言数学题(无思维链标注)上训练,polyGRPO使基线模型(Qwen2.5-7B-Instruct)在四个英文推理测试集上绝对提升6.72%,在多语言基准上提升6.89%。尤为突出的是,它是唯一在英文常识推理任务(4.9%)上超越基线的方法,尽管训练仅使用数学数据,体现强大跨任务泛化能力。进一步分析显示,将语言视为潜在变量可扩展模型的潜在推理空间,带来一致且可推广的性能提升。

原文摘要 · Abstract (English)

As LLMs reduce English-centric bias, a surprising trend emerges: non-English responses sometimes outperform English on reasoning tasks. We hypothesize that language functions as a latent variable that structurally modulates the model's internal inference pathways, rather than merely serving as an output medium. To test this, we conducted a Polyglot Thinking Experiment, in which models were prompted to solve identical problems under language-constrained and language-unconstrained conditions. Results show that non-English responses often achieve higher accuracy, and the best performance frequently occur when language is unconstrained, suggesting that multilinguality broadens the model's latent reasoning space. Based on this insight, we propose polyGRPO (Polyglot Group Relative Policy Optimization), an RL framework that treats language variation as an implicit exploration signal. It generates polyglot preference data online under language-constrained and unconstrained conditions, optimizing the policy with respect to both answer accuracy and reasoning structure. Trained on only 18.1K multilingual math problems without chain-of-thought annotations, polyGRPO improves the base model (Qwen2.5-7B-Instruct) by 6.72% absolute accuracy on four English reasoning testset and 6.89% in their multilingual benchmark. Remarkably, it is the only method that surpasses the base LLM on English commonsense reasoning task (4.9%), despite being trained solely on math data-highlighting its strong cross-task generalization. Further analysis reveals that treating language as a latent variable expands the model's latent reasoning space, yielding consistent and generalizable improvements in reasoning performance.

推理优化多语言强化学习通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。