arXiv:2508.08204cs.CLcs.AI2025-08被引 1

评估大模型推理时不确定性与人类认知的匹配度,发现部分指标既贴近人感又具备校准能力。

Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models

  • 测试多种推理时不确定性度量方法,对比其与人类集体不确定性的匹配程度。
  • 多个指标显示与人类不确定性高度一致,且在正确性相关性和分布分析中表现校准良好。
  • 适用于需提升人机信任、优化交互体验的LLM应用,尤其关注可解释性与可控性。

近期研究关注大语言模型的不确定性校准,以增强模型控制并调节用户信任。推理时不确定性可为模型或外部控制模块提供实时信号,对实际提升人机交互体验至关重要。尽管已有大量工作探讨模型校准,但较少研究考察模型不确定性与人类不确定性之间的对齐程度。本文评估了多种推理时不确定性度量方法,结合现有指标与新变体,检验其与人类群体级不确定性及传统校准概念的契合度。结果表明,诸多度量指标展现出与人类不确定性较强的对齐性,即使在未对齐人类答案偏好时亦然。对于表现优异的指标,我们观察到其在正确性相关性和分布分析方面具有中等到强的校准证据。

原文摘要 · Abstract (English)

There has been much recent interest in evaluating large language models for uncertainty calibration to facilitate model control and modulate user trust. Inference time uncertainty, which may provide a real-time signal to the model or external control modules, is particularly important for applying these concepts to improve LLM-user experience in practice. While many of the existing papers consider model calibration, comparatively little work has sought to evaluate how closely model uncertainty aligns to human uncertainty. In this work, we evaluate a collection of inference-time uncertainty measures, using both established metrics and novel variations, to determine how closely they align with both human group-level uncertainty and traditional notions of model calibration. We find that numerous measures show evidence of strong alignment to human uncertainty, even despite the lack of alignment to human answer preference. For those successful metrics, we find moderate to strong evidence of model calibration in terms of both correctness correlation and distributional analysis.

不确定性人机对齐校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。