将语言模型概率转化为可验证的语义不确定性度量,提升决策可信度。
Calibrating Semantic Uncertainty from Observable Language-Model Probabilities

- 构建语义映射,将文本概率转换为可检验的状态后验
- 在真实市场数据上实现有效不确定性覆盖,且对改写保持稳定
- 适合需可靠推理的科研与金融决策场景
随着生成式AI进入科学与专业领域,其不确定性需定义在影响推理和决策的关键状态上。语言模型仅给出词级概率,而实际应用需要诊断、假设或运行状态等有意义状态的不确定性。本文提出一种语义映射:预先设定、可检验的桥梁,将语言响应的概率映射到有限状态的后验分布。语言分布无需受限;通过保留校准将其连接至参考后验。推导了后验误差边界及存在性、条件唯一性、表示稳定性与稳定逆恢复的条件。该区分至关重要,因语言概率依赖提示措辞,而目标后验应不随信息等价的重述改变。实验使用美联储经济与金融系列整理的专业市场文本,结合具有精确后验的受控模拟。在两个训练好的语言模型中,语言衍生概率优于人工数值置信度,能准确恢复保留后验,具备有效的不确定性覆盖率,在改写下保持稳定,并对证据变化作出合理响应。提示工程优化的是依赖措辞的输出,而稳健的科学应用要求与任务相关含义的可验证稳定性。所提映射将生成系统中的语义不确定性转化为可识别、可检验的统计问题,当接受条件满足时,可提供可审计的后验估计。
原文摘要 · Abstract (English)
As generative artificial intelligence enters scientific and professional work, its uncertainty must be defined on the states that matter for inference and decision-making. Language models assign probabilities to words, whereas applications require uncertainty over meaningful states such as diagnoses, hypotheses or operational conditions. We introduce a \emph{semantic map}: a prespecified, testable bridge from probabilities over verbal responses to a posterior over declared finite states. The language distribution remains unrestricted; held-out calibration connects it to a reference posterior. We derive posterior-error bounds and conditions for existence, conditional uniqueness, presentation stability and stable inverse recovery. This distinction matters because language probabilities depend on prompt wording, while the target posterior should not change under information-equivalent rewording. Experiments use professional market text compiled from Federal Reserve economic and financial series, together with controlled simulations having exact posteriors. Across two fitted language models, language-derived probabilities outperform printed numerical confidence, recover held-out posteriors with valid uncertainty coverage, remain largely stable under paraphrase and respond appropriately to altered evidence. \textbf{Prompt engineering optimises a wording-dependent response; robust scientific use requires validated stability of application-relevant meaning.} The proposed map turns semantic uncertainty in generative systems into an identifiable and testable statistical measurement problem and, when its acceptance conditions hold, yields an auditable posterior estimate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。