提出严格满足概率准确性的智能体不确定性评估方法
Proper Scoring Rules for Agentic Uncertainty Quantification
- 设计轨迹正则评分TPS,精准捕捉每一步的成功概率序列
- 实验证明重校准可大幅提升TPS但对排名指标影响小
- 适合需要真实概率输出的高可靠性决策场景
语言模型智能体在推理过程中不断输出不确定性信号,但现有评估方法常混淆排序有效性与概率真实性。当前常用指标如AUROC、AUPRC、风险覆盖、轨迹ECE及标量化的轨迹评分,仅衡量区分度、分组校准或压缩结果,无法严格激发完整的前缀条件成功概率轨迹 $q_t = P^π(Y=1 | H_t)$。本文基于预序正则评分,提出轨迹正则评分(TPS),为任意步骤不确定性信号转化为最终成功概率的轨迹提供预测器无关的严格正则评分族。理论证明,在完整观测下,TPS能严格激发成功概率过程。针对实际中轨迹因停止而部分可观测的情况,通过投影完整数据得分至可观测前缀,得到精确的 $q_Z$-加权简化评分,并给出 $q_Z$ 未估计时的可计算近似。进一步分析表明,常见轨迹评估器针对的是比完整前缀条件概率更弱的对象:轨迹ECE对分辨率不敏感,标量化的轨迹Brier仅激发压缩标量而非完整轨迹。在StrategyQA、Tau2-Bench、HotpotQA和WebShop上的实验显示,这些理论差异具有实际可操作性:概率重校准可显著提升TPS值,而排名指标几乎不变;且在缺失完整轨迹时,近似评分可能导致与完整评估不同的结论。
原文摘要 · Abstract (English)
Language-model agents increasingly emit uncertainty signals throughout a trajectory, but existing agentic UQ evaluations often conflate ranking usefulness with probabilistic truthfulness. AUROC, AUPRC, risk-coverage, Trajectory ECE, and scalarized trajectory scores evaluate discrimination, binwise calibration, or collapsed summaries, but do not strictly elicit the full prefix-conditioned success-probability trace $q_t = P^π(Y=1 | H_t)$. Building on prequential proper scoring, we introduce the Trajectory Proper Score (TPS), a predictor-agnostic family of strictly proper trajectory-level scoring rules for any per-step uncertainty signal calibrated into a probability of eventual success. We prove that TPS strictly elicits the success-probability process under complete observation, within the chosen score family and weight schedule. We extend the construction to administratively censored trajectories by projecting the complete-data score onto the observable stopped prefix, yielding an exact $q_Z$-weighted reduced score and a tractable approximation when $q_Z$ is unestimated. We further show that common trajectory evaluators target weaker objects than the full prefix-conditioned probability process: Trajectory ECE is resolution-blind, while scalarized Trajectory Brier elicits only the collapsed scalar, not the full trace. Experiments on StrategyQA, Tau2-Bench, HotpotQA, and WebShop show that these theoretical distinctions are operationally visible: probability recalibration can substantially change TPS while leaving rank metrics nearly unchanged, and the tractable censored approximation can change the verdict relative to complete-only evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。