多专家协作+一致性验证,让医疗问答模型更懂自己知道多少。
Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQA
- 四个专科专家独立答题,通过两阶段自检生成可信度评分
- 在医学多选题上,误差率降低74.4%,准确率达59.2%
- 适合需要可信置信度的临床AI系统开发与评估
错误的置信度会阻碍AI在临床中的应用。本文提出一种多代理框架,结合领域专精代理与两阶段验证(Wu et al., 2024)及S-score加权融合,提升医学多选题回答的校准性与区分度。四名专家(呼吸、心血管、神经、消化)使用Qwen2.5-7B-Instruct独立生成诊断,每项诊断经过两阶段自检以衡量内部一致性并生成专家置信度(S-score)。S-score驱动加权融合策略,决定最终答案并校准报告置信度。在MedQA-USMLE和MedMCQA的高分歧子集(100和250题)上评估。在MedQA-250上,系统实现ECE = 0.091(比单专家基线降低74.4%),AUROC = 0.630(+0.056),准确率为59.2%。所有设置下校准性能提升49%-74%。消融分析显示,两阶段验证主导ECE下降,多代理推理提升AUROC,表明一致性检查与集成聚合分别应对不同类型的置信度失效。该置信信号是否足以支持临床拒答决策,尚待未来研究。
原文摘要 · Abstract (English)
Miscalibrated confidence scores are a practical obstacle to deploying AI in clinical settings. A model that is always overconfident offers no useful signal for deferral. We present a multi-agent framework that combines domain-specific specialist agents with Two-Phase Verification (Wu et al., 2024) and S-Score Weighted Fusion to improve both calibration and discrimination in medical multiple-choice question answering. Four specialist agents (respiratory, cardiology, neurology, gastroenterology) generate independent diagnoses using Qwen2.5-7B-Instruct. Each diagnosis undergoes a two-phase self-verification process that measures internal consistency and produces a Specialist Confidence Score (S-score). The S-scores drive a weighted fusion strategy that selects the final answer and calibrates the reported confidence. We evaluate on high-disagreement subsets of MedQA-USMLE and MedMCQA (100 and 250 questions). All results are specific to this filtered regime. On MedQA-250, the full system achieves ECE = 0.091 (74.4% reduction over the single-specialist baseline) and AUROC = 0.630 (+0.056) at 59.2% accuracy. Calibration gains of 49-74% hold across all four settings. Ablation analysis reveals that Two-Phase Verification drives ECE reduction while multi-agent reasoning drives AUROC improvement, suggesting that consistency checking and ensemble aggregation address different failure modes of LLM uncertainty. Whether the resulting confidence signal is sufficient to support clinical deferral decisions in practice remains a direction for future investigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。