arXiv:2603.02798cs.AIcs.CL2026-03被引 3

用临床指南构建可信赖的AI诊断验证系统

Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification

  • 基于专家指南构建多阶段决策对齐评估机制
  • 在三种疾病上提升12%的判别能力,降低50%误差
  • 适合医疗AI部署中的可靠性验证场景

随着大模型代理被用于高风险决策(如临床诊断),其决策可靠性验证变得至关重要。现有验证方法因缺乏领域知识和校准不足而表现不佳。为此,我们提出GLEAN框架,通过将专家制定的诊疗指南转化为轨迹感知的校准正确性信号,评估代理每一步是否符合指南,并聚合多指南评分生成代理特征,沿决策路径累积并用贝叶斯逻辑回归校准为正确概率。同时,估计的不确定性触发主动验证:对不确定案例通过扩展指南覆盖范围和差异检查主动收集额外证据。我们在MIMIC-IV数据集上的三类疾病临床诊断任务中验证了GLEAN,AUROC相比最优基线提升12%,Brier得分降低50%,证明其在判别力与校准性上的有效性。临床专家研究也认可其实际应用价值。

原文摘要 · Abstract (English)

As LLM-powered agents have been used for high-stakes decision-making, such as clinical diagnosis, it becomes critical to develop reliable verification of their decisions to facilitate trustworthy deployment. Yet, existing verifiers usually underperform owing to a lack of domain knowledge and limited calibration. To address this, we establish GLEAN, an agent verification framework with Guideline-grounded Evidence Accumulation that compiles expert-curated protocols into trajectory-informed, well-calibrated correctness signals. GLEAN evaluates the step-wise alignment with domain guidelines and aggregates multi-guideline ratings into surrogate features, which are accumulated along the trajectory and calibrated into correctness probabilities using Bayesian logistic regression. Moreover, the estimated uncertainty triggers active verification, which selectively collects additional evidence for uncertain cases via expanding guideline coverage and performing differential checks. We empirically validate GLEAN with agentic clinical diagnosis across three diseases from the MIMIC-IV dataset, surpassing the best baseline by 12% in AUROC and 50% in Brier score reduction, which confirms the effectiveness in both discrimination and calibration. In addition, the expert study with clinicians recognizes GLEAN's utility in practice.

AI医疗验证框架大模型可靠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。