arXiv:2604.00025cs.CLcs.AI2026-04被引 4

让大模型说短话,反而比小模型更准,还能省算力。

Brevity Constraints Reverse Performance Hierarchies in Language Models

论文配图:Brevity Constraints Reverse Performance Hierarchies in Language Models
图 1 · 摘自论文原文
  • 用简洁指令约束大模型,避免啰嗦导致错误。
  • 数学和科学题上大模型准确率反超小模型7.7~15.9个百分点。
  • 适合关注推理能力与高效部署的研究者和工程师。

标准评测发现,在五个数据集共1,485个问题中,有7.7%的问题上,参数量大10-100倍的模型反而比小模型低28.4个百分点。通过系统评估31个模型(0.5B-405B参数),我们识别出原因是大模型自发产生冗长回应,引发过拟合错误。因果干预实验表明,这是可纠正的提示设计问题,而非能力缺陷。强制大模型输出简短回答,准确率提升26个百分点,性能差距缩小至原来的三分之一。尤其在数学推理和科学知识任务上,大模型优势逆转,反超小模型7.7-15.9个百分点。这证明大模型具备隐藏的强推理能力,被通用提示掩蔽。三重污染测试验证结果可靠,逆向缩放现象在全参数范围内连续存在,最优规模因数据集而异(0.5B–3.0B)。结论:提升大模型表现需基于规模感知的提示工程,而非统一评测。此策略同时提升精度并降低计算成本。

原文摘要 · Abstract (English)

Standard evaluation protocols reveal a counterintuitive phenomenon: on 7.7% of benchmark problems spanning five datasets, larger language models underperform smaller ones by 28.4 percentage points despite 10-100x more parameters. Through systematic evaluation of 31 models (0.5B-405B parameters) across 1,485 problems, we identify the mechanism as spontaneous scale-dependent verbosity that introduces errors through overelaboration. Causal intervention experiments demonstrate this reflects correctable prompt design rather than fundamental capability limitations. Constraining large models to produce brief responses improves accuracy by 26 percentage points and reduces performance gaps by up to two-thirds. Most critically, brevity constraints completely reverse performance hierarchies on mathematical reasoning and scientific knowledge benchmarks, with large models achieving 7.7-15.9 percentage point advantages over small models -- direct inversions of the original gaps. These reversals prove large models possess superior latent capabilities that universal prompting masks. We validate findings through three independent contamination tests and demonstrate inverse scaling operates continuously across the full parameter spectrum, with dataset-specific optimal scales ranging from 0.5B to 3.0B parameters. Our results establish that maximizing large model performance requires scale-aware prompt engineering rather than universal evaluation protocols, with immediate implications for deployment: prompt adaptation simultaneously improves accuracy and reduces computational costs.

大模型提示工程推理优化逆向缩放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。