arXiv:2604.08563cs.CLcs.AI2026-04被引 1

调温度能提升大模型推理表现,零样本提示在中温下最有效。

Temperature-Dependent Performance of Prompting Strategies in Extended Reasoning Large Language Models

论文配图:Temperature-Dependent Performance of Prompting Strategies in Extended Reasoning Large Language Models
图 1 · 摘自论文原文
  • 测试四种温度下零样本与思维链提示的效果差异。
  • 零样本在0.4和0.7温度下达59%准确率,思维链在极端温度表现更好。
  • 温度越高,扩展推理增益越明显,最高达14.3倍。

扩展推理模型通过显式测试时计算实现了大语言模型能力的飞跃,但采样温度与提示策略的最佳组合仍不明确。本文在Grok-4.1上,针对AMO-Bench(39道国际数学奥林匹克级难题)的扩展推理能力,系统评估了零样本与思维链提示在四个温度设置(0.0、0.4、0.7、1.0)下的表现。结果发现,零样本提示在中等温度下性能最佳,0.4和0.7时准确率达59%;而思维链提示在温度极值下表现更优。尤为关键的是,扩展推理的增益随温度升高从6倍提升至14.3倍。这表明温度应与提示策略协同优化,挑战了当前普遍采用T=0进行推理的做法。

原文摘要 · Abstract (English)

Extended reasoning models represent a transformative shift in Large Language Model (LLM) capabilities by enabling explicit test-time computation for complex problem solving. However, the optimal configuration of sampling temperature and prompting strategy for these systems remains largely underexplored. We systematically evaluate chain-of-thought and zero-shot prompting across four temperature settings (0.0, 0.4, 0.7, and 1.0) using Grok-4.1 with extended reasoning on 39 mathematical problems from AMO-Bench, a challenging International Mathematical Olympiad-level benchmark. We find that zero-shot prompting achieves peak performance at moderate temperatures, reaching 59% accuracy at T=0.4 and T=0.7, while chain-of-thought prompting performs best at the temperature extremes. Most notably, the benefit of extended reasoning increases from 6x at T=0.0 to 14.3x at T=1.0. These results suggest that temperature should be optimized jointly with prompting strategy, challenging the common practice of using T=0 for reasoning tasks.

大模型推理温度调控提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。