调温度能影响大模型评分的稳定性和探索性,需根据任务灵活设置。
The Necessity of Setting Temperature in LLM-as-a-Judge
- 通过系统实验发现温度影响评分一致性与格式错误率。
- 高温使模型暴露不确定性,尤其在模糊场景中更易出错。
- 复杂或模糊任务适合高温探索,稳定任务宜用低温保证一致。
将大语言模型(LLM)用作评估模型输出的裁判,已成为自动化评估的重要范式。然而,解码温度的选择仍主要依赖经验,缺乏系统的实证研究。为此,我们系统研究了温度对不同LLM裁判模型、提示策略和评估范式下判断行为的影响。结果表明,较高温度通常降低判断一致性并增加格式错误,同时揭示了低温度下被压制的潜在不确定性,尤其在模糊案例中更为明显。进一步分析表明,较高温度可作为探索机制,在复杂或不确定的评估场景中可能提升裁判表现。总体而言,低温度更适合强调稳定性和可复现性的任务,而高温度则适用于存在显著模糊性或复杂性的场景,此时对裁判决策空间的探索更有益。这些发现表明,在LLM-as-a-judge系统中,温度不应视为固定超参数,而应作为可控的、任务相关的决策变量,权衡可靠性与探索性。
原文摘要 · Abstract (English)
Using large language models (LLMs) as judges for evaluating model outputs has emerged as an important paradigm for automated evaluation. However, the choice of decoding temperature in LLM-as-a-judge settings is still largely chosen empirically, with limited systematic evidence on its impact. To address this gap, we conduct a systematic study of how temperature affects judgment behavior across different LLM judge models, prompting strategies, and evaluation paradigms. Our results show that higher temperatures generally decrease judgment consistency and increase formatting errors, while also exposing latent uncertainty that tends to remain suppressed under low-temperature decoding, particularly in ambiguous cases. Further analysis suggests that higher temperatures can serve as an exploratory mechanism and may improve judging performance in complex or uncertain evaluation scenarios. Overall, low-temperature settings are better suited to tasks that prioritize stability and reproducibility, whereas higher-temperature settings are more appropriate for scenarios involving substantial ambiguity or complexity, where exploration of the judge's decision space is beneficial. These findings suggest that, in LLM-as-a-judge systems, temperature should be treated not as a fixed hyperparameter, but as a controllable, task-dependent design choice that mediates the trade-off between reliability and exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。