arXiv:2603.19118cs.AIcs.CL2026-03被引 1

研究推理模型不确定性估计在采样增多时的表现,发现组合信号效果更好。

How Uncertainty Estimation Scales with Sampling in Reasoning Models

  • 用口头置信度和自洽性做黑箱不确定性估计
  • 两样本组合即提升效果,大幅优于单一信号
  • 数学任务中效果更优,且信号互补性强

不确定性估计对部署推理语言模型至关重要,但在扩展链式思维推理中仍理解不足。本文通过口头置信度与自洽性,研究并行采样在三种推理模型、17个跨数学、STEM和人文学科任务中的表现。结果表明,自洽性和口头置信度均随采样增长而提升,但自洽性初始区分度较低且在中等采样下落后于口头置信度。大部分不确定性增益来自信号组合:仅用两个样本,混合估计算法平均提升AUROC达+12,已超越任一单独信号,即使扩大采样预算后收益也趋于饱和。该效应具有领域依赖性:在数学任务中,基于RLVR风格后训练的模型具备更高不确定性质量,表现出更强互补性与更快的缩放速度。

原文摘要 · Abstract (English)

Uncertainty estimation is critical for deploying reasoning language models, yet remains poorly understood under extended chain-of-thought reasoning. We study parallel sampling as a fully black-box approach using verbalized confidence and self-consistency. Across three reasoning models and 17 tasks spanning mathematics, STEM, and humanities, we characterize how these signals scale. Both self-consistency and verbalized confidence scale in reasoning models, but self-consistency exhibits lower initial discrimination and lags behind verbalized confidence under moderate sampling. Most uncertainty gains, however, arise from signal combination: with just two samples, a hybrid estimator improves AUROC by up to $+12$ on average and already outperforms either signal alone even when scaled to much larger budgets, after which returns diminish. These effects are domain-dependent: in mathematics, the native domain of RLVR-style post-training, reasoning models achieve higher uncertainty quality and exhibit both stronger complementarity and faster scaling than in STEM or humanities.

不确定性估计推理模型自洽性采样效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。