不同语言需不同词数表达相同内容,输出长度限制会扭曲多语言推理能力评估结果。
Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap

- 将输出令牌上限设为独立变量,测试其对多语言推理差距的影响。
- 在紧约束下,语言差距最大波动达57分,长度归一化可移动38.9分。
- 应报告跨预算范围的准确率变化,而非单一数值,适合评估多语言模型的研究者。
多语言评估通常在固定输出令牌数下进行,但不同语言表达相同内容所需令牌数不同,导致该上限成为隐藏的实验变量。本文在Qwen3-8B和Llama-3.1-8B-Instruct上,针对德语、泰语、斯瓦希里语(MGSM)测试了原生与翻译之间的差距是否由令牌预算造成,涵盖四种提示策略。结果显示,测量差距在不同预算下最高波动57分;当预算紧缩时,长度归一化可使结果移动最多38.9分;在严格限制下,归一化甚至能反转策略表现排序。研究前瞻性冻结了三组Qwen峰值表现及1024时接近零的值,并在54万次独立硬性截断解码中验证,六组霍尔姆校正检验均拒绝原假设。在1024预算下仍无法拒绝原假设,因原生准确率已饱和;此后残差差异实为策略性能差,非推理缺陷。交叉拟合的泰语词汇扩展在冻结预算下贡献0.0分,在19%轨迹仍截断时提升4.9分。第三组冻结实验仅改变公布预算(128 vs 2048),泰语原生准确率变动5.1分,说明准确率不仅取决于强制上限。基于一次长上限运行计算的正确生成时间恒等式,与预设三个MGSM峰值匹配至0.65分,且在另三个基准的探索性分析中,对保留样本预测误差仅0.92分,五处位置精准定位峰值。建议将输出上限视为独立变量,报告跨预算区间的准确率表现,而非单一预算下的数值。
原文摘要 · Abstract (English)
Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable. We test whether the native-vs-translate gap on MGSM (German, Thai, Swahili) is a token-budget artifact for Qwen3-8B and Llama-3.1-8B-Instruct under four prompting strategies. The measured gap swings by up to 57 points across budgets, length normalization moves it by up to 38.9 points where the cap binds, and at tight caps normalization can reverse which strategy scores higher. We prospectively froze the sweep's three Qwen peaks and its near-zero value at 1024 and evaluated them on 540,000 independently hard-capped decodes: a second frozen family of six Holm-corrected tests rejects every null. The frozen test at $B^*=1024$ still fails to reject because native accuracy has already saturated there; above saturation, the residual difference is a strategy-performance gap, not an identified reasoning deficit. The same truncation channel prices a cost-ordered adaptation ladder: a cross-fitted Thai vocabulary extension closes 0.0 points of the gap at the frozen budget and 4.9 points where 19% of traces still truncate. A third frozen family varies only the announced budget at a fixed enforced cap; announcing 128 rather than 2048 tokens moves Thai native accuracy by 5.1 points, so accuracy is not a function of the enforced cap alone. A correct-emission timing identity computed from one long-cap run matches the three pre-specified MGSM peaks to 0.65 points and, in an exploratory Qwen-only analysis of three further benchmarks, tracks held-out items to 0.92 points, locating the peak exactly in five of seven cells. Treat the output cap as an independent variable and report accuracy across the budget regime, not at a single budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。