arXiv:2606.11512cs.CL2026-06被引 1

让AI的不确定表达更真实,通过反复生成结果来校准其说法。

SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment

论文配图:SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment
图 1 · 摘自论文原文
  • 基于多次生成结果,构建与答案相关的不确定性几何结构
  • 在事实、数学等任务中降低过自信,提升不确定性排序准确率
  • 适合需要可靠可信度判断的AI系统开发者

大型语言模型越来越多地通过自然语言表达不确定性,但这些表达往往无法反映模型的真实采样行为。本文将口语化不确定性对齐视为分布校准问题:针对特定提示的适当不确定性目标应从重复模型输出中估计,而非单一响应。仅使用群体生成仍不足,因目标需提供有效训练信号。现有目标部分满足要求。本文提出SAGE(语义-答案引导熵),一种群体级不确定性目标,通过采样响应构建答案条件下的不确定性几何结构。SAGE保持类别、数值和符号答案的区分性,同时维持平滑且尺度不变的校准信号。进一步通过组不确定性偏好优化(GUPO)应用该目标,这是一种监督口语化不确定性表达而非完整响应的训练框架。跨事实、数学及多选推理任务的实验表明,其显著改善不确定性排序,降低校准误差,减少过自信现象。

原文摘要 · Abstract (English)

Large language models increasingly express uncertainty through natural-language statements, yet these expressions often fail to reflect the model's sampled behavior. We study verbal uncertainty alignment as a distributional calibration problem: the appropriate uncertainty target for a prompt should be estimated from repeated model outputs rather than from an isolated response. However, group rollouts alone are insufficient, since the resulting target must provide a useful training signal. Existing targets only partially satisfy this requirement. We propose SAGE, Semantic-Answer Guided Entropy, a group-level uncertainty target that constructs an answer-conditioned uncertainty geometry over sampled responses. SAGE preserves categorical, numeric, and symbolic answer distinctions while maintaining a smooth and scale-preserving calibration signal. We further apply this target through Group-Uncertainty Preference Optimization, or GUPO, an uncertainty-channel training framework that supervises verbal uncertainty expressions rather than the full response. Experiments across factual, mathematical, and multiple-choice reasoning tasks show improved uncertainty ranking, lower calibration error, and reduced overconfidence.

不确定性建模语言模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。