arXiv:2508.15050cs.AIcs.CL2025-08被引 4

推理越深越不靠谱,找资料比死磕更有效。

Don't Think Twice! Over-Reasoning Impairs Confidence Calibration

  • 用检索增强生成替代深度推理,提升信心校准
  • 增加思考时间反而让模型更自大,准确率降至48.7%
  • 适合需要可靠置信度评估的AI问答系统开发者

将大语言模型作为问答工具时,信心校准能力至关重要。我们系统评估了推理能力与计算预算对信心判断准确性的影响,使用ClimateX数据集(Lacombe等,2023)并扩展至人类与行星健康领域。关键发现挑战了“测试时缩放”范式:尽管最新推理模型在评估专家信心方面达到48.7%准确率,但增加推理预算反而持续损害校准效果。更长的思考时间导致系统性过度自信,计算投入超过适度水平后出现收益递减甚至负收益。相比之下,检索增强生成显著优于纯推理,通过获取相关证据实现89.3%的准确率。结果表明,对于知识密集型任务,信息获取能力而非推理深度或推理预算,可能是信心校准的关键瓶颈。

原文摘要 · Abstract (English)

Large Language Models deployed as question answering tools require robust calibration to avoid overconfidence. We systematically evaluate how reasoning capabilities and budget affect confidence assessment accuracy, using the ClimateX dataset (Lacombe et al., 2023) and expanding it to human and planetary health. Our key finding challenges the "test-time scaling" paradigm: while recent reasoning LLMs achieve 48.7% accuracy in assessing expert confidence, increasing reasoning budgets consistently impairs rather than improves calibration. Extended reasoning leads to systematic overconfidence that worsens with longer thinking budgets, producing diminishing and negative returns beyond modest computational investments. Conversely, search-augmented generation dramatically outperforms pure reasoning, achieving 89.3% accuracy by retrieving relevant evidence. Our results suggest that information access, rather than reasoning depth or inference budget, may be the critical bottleneck for improved confidence calibration of knowledge-intensive tasks.

LLM校准推理质量检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。