arXiv:2509.24202cs.CLcs.AI2025-09被引 5

让大模型用自然语言表达不确定,更可信更高效。

Can Large Language Models Express Uncertainty Like Human?

  • 用语气词表达不确定,无需复杂计算
  • 新数据集+轻量映射器,低成本转信心分数
  • 提示工程可让模型表现接近人类,适合高风险场景

大语言模型在高风险场景中广泛应用,但过度自信的回应可能误导用户。可靠的置信度估计能提升信任与任务准确率。然而现有方法存在实际障碍:逻辑值常被隐藏,多采样计算成本高,口头数值不确定性(如0-100分)偏离自然交流。本文重新审视语言置信度(LC),即模型通过模糊表达(如‘可能’、‘或许’)传递不确定性的轻量级、以人为本方式。为此,我们(1)发布首个多样且大规模的模糊表达数据集,含人工标注置信度;(2)提出轻量级映射器,以近乎零成本将模糊语句转为置信分数。基于此,(3)首次系统研究现代大模型在问答基准上的语言置信表现,发现多数模型表达不可靠,但精心设计的提示可实现竞争性校准与区分能力。最后,(4)引入微调框架进一步提升表达可靠性。综合来看,本工作将语言置信定位为一种可扩展、高效且符合人类认知的不确定性估计方案,并呼吁深入探索这一有前景却未被充分研究的方向。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in high-stakes settings, where overconfident responses can mislead users. Reliable confidence estimation has been shown to enhance trust and task accuracy. Yet existing methods face practical barriers: logits are often hidden, multi-sampling is computationally expensive, and verbalized numerical uncertainty (e.g., giving a 0-100 score) deviates from natural communication. We revisit linguistic confidence (LC), where models express uncertainty through hedging language (e.g., probably, might), offering a lightweight and human-centered alternative. To advance this direction, we (1) release the first diverse, large-scale dataset of hedging expressions with human-annotated confidence scores, and (2) propose a lightweight mapper that converts hedges into confidence scores at near-zero cost. Building on these resources, we (3) conduct the first systematic study of LC across modern LLMs and QA benchmarks, revealing that while most LLMs underperform in expressing reliable LC, carefully designed prompting achieves competitive calibration and discriminability. Finally, we (4) introduce a fine-tuning framework that further improves LC reliability. Taken together, our work positions linguistic confidence as a scalable, efficient, and human-aligned approach to LLM uncertainty estimation, and calls for deeper exploration of this promising yet underexplored direction.

不确定性语言模型可信生成提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。