arXiv:2504.14047cs.AI2025-04被引 20

推理模型需内化思维逻辑,仅靠增加推理计算无法突破性能天花板。

Rethinking Inference-Time Scaling: Efficiency Limits and Linguistic Signals

  • 发现非推理优化模型存在不可逾越的推理下限,即使加倍推理算力也难提升
  • 简单多数投票优于复杂修正框架,表明方法越简单效果越好
  • 正确回答更简洁、少用模糊词,可作零算力质量自检信号

研究深入分析了在挑战性推理任务上,通用模型与推理优化模型在推理时计算(ITC)扩展中的表现。尽管已有研究认为推理算力可替代模型参数扩展,但本文揭示了根本性限制:非推理优化模型存在‘推理下限’,无论投入多少推理算力都无法突破。实验显示,通用模型即使使用十倍于推理优化模型的推理算力,仍无法达到其准确率。在推理模型内部,复杂推理策略常带来边际收益递减,而简单的多数投票始终表现更优。关键发现是:正确答案显著更简洁,且减少使用模糊表达和思考标记,这些语言特征可作为无需额外算力的质量判断代理,为构建高效自诊断推理系统提供新路径。

原文摘要 · Abstract (English)

There is intense interest in investigating how inference time compute (ITC) (e.g. repeated sampling, refinements, etc) can improve large language model (LLM) capabilities. While breakthroughs like DeepSeek-R1 highlight the power of reinforcement learning for reasoning, the interaction between ITC and reasoning-optimized weights remains poorly understood. This work conducts a comprehensive analysis of inference-time scaling methods for both reasoning and non-reasoning models on challenging reasoning tasks. While prior work suggests that scaling test-time compute can optimally substitute for model parameter scaling, we identify a fundamental limit to this compute-equivalence, the reasoning floor, a performance plateau that non-reasoning models cannot escape, no matter how much inference compute is spent. We demonstrate that general-purpose models fail to match the accuracy of reasoning-optimized models even with an order of magnitude more inference compute, suggesting that internalizing reasoning protocols is a prerequisite for effective test-time scaling. Within reasoning models, we find that the complexity of the scaling method often yields diminishing returns; simple majority voting consistently outperforms sophisticated sequential revision and mixture-of-agents frameworks. Crucially, we identify a Linguistic Signal of Correctness - correct responses are significantly more concise and exhibit a lower density of hedging and thinking markers. We demonstrate that these intrinsic linguistic features can serve as zero-compute proxies for response quality, providing a pathway to more efficient, self-diagnostic reasoning agents.

大模型推理语言信号效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。