arXiv:2608.20717cs.AI2026-08中稿 · PRICAI 2026

用狄利克雷证据聚合提升数学推理中语言自信的校准精度

DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

论文配图:DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning
图 1 · 摘自论文原文
  • 将多轮提示下的自信值转化为软证据,建模答案不确定性
  • 在GSM8K等数据集上校准误差降低23%,同时保持高准确率
  • 适合需要可信自信度评估的数学推理系统开发者

可靠的置信度估计对大语言模型在数学推理中的应用至关重要,但黑箱式语言自信难以校准。当同一问题使用多种引导性提示查询时,得到的答案-自信观测值包含有用不确定性信息,但其数值尺度可能随提示级别、模型和数据集变化。现有黑箱不确定性方法依赖答案一致性、样本一致性或熵,仅描述输出变异,未建模自报自信的数值意义。直接平均或启发式聚合无法学习提示与任务相关的偏差。本文提出DirEAG,一种狄利克雷证据聚合方法,将每条诱发的答案-自信观测值转换为生成候选答案及额外空状态上的校准软证据,使模型能表示所有候选均不正确的情况。在Qwen、Mistral、Gemma模型上,于GSM8K、SVAMP和GSM-Hard数据集的实验表明,相比直接平均和启发式聚合,DirEAG通常实现更优校准,同时保持竞争力的答案选择能力。消融实验进一步揭示证据聚合与最终二元校准分别解决校准问题的不同方面。

原文摘要 · Abstract (English)

Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confidence-steering prompts, the resulting answer-confidence observations contain useful uncertainty information, yet their scales may shift across steering levels, models, and datasets. Existing black-box uncertainty methods often rely on answer agreement, sample consistency, or entropy, which describe output variation but do not model the numerical meaning of self-reported confidence. Conversely, direct averaging or heuristic aggregation of elicited confidence cannot learn prompt- and task-dependent bias. We propose DirEAG, a Dirichlet Evidence Aggregation method that converts each elicited answer-confidence observation into calibrated soft evidence over generated candidate answers and an additional null state, allowing the model to represent cases where none of the candidates is correct. Experiments on GSM8K, SVAMP, and GSM-Hard with Qwen, Mistral, and Gemma models show that, compared with direct confidence averaging and heuristic confidence-steering aggregation, DirEAG often achieves better calibration while maintaining competitive answer selection. Ablations further reveal that evidence aggregation and final binary calibration address distinct parts of the calibration problem.

数学推理置信度校准证据聚合大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。