分离模型能力与置信度,提升医疗视觉问答的可信度。
Separating Capability from Confidence: Grounded Dual-State Calibration for GRPO-Trained Medical Vision-Language Models

- 将置信度估计与答案生成解耦,利用内部状态评估可靠性。
- 在两个基准上提升正确性区分度和校准性能,保持答案准确率。
- 引入反事实置信度接地指标,验证置信度是否基于视觉证据。
医疗视觉语言模型(VLMs)需要反映答案正确性与患者特定视觉证据的置信度。近期基于GRPO的方法联合优化表述性置信度与答案生成,但这种联合优化可能干扰答案学习,并使置信度趋向二值化。表述性置信度也无法明确评估视觉支持。因此,本文分离能力学习与置信度估计,提出 extbf{DualRead}。DualRead基于关键回答时刻的行动者内部状态可读取可靠性这一洞察,冻结已训练的GRPO行动者,结合回答前可解性与回答后对生成答案及其视觉支持的评估。为进一步检验置信度是否反映视觉基础,引入 extbf{反事实置信度接地AUC}(CCG-AUC),衡量当真实图像替换导致行动者从正确变为错误时,置信度是否下降。在两个VLM主干网络及分布内/外医疗VQA基准上,DualRead在保持答案准确率的同时,提升了正确性区分度与校准性能。CCG-AUC揭示了置信度是否响应与答案相关的视觉证据,而非主要依赖非视觉线索。
原文摘要 · Abstract (English)
Medical vision-language models (VLMs) require confidence that reflects both answer correctness and patient-specific visual evidence. Recent GRPO-based methods optimize verbalized confidence together with answer generation. However, this joint optimization may interfere with answer learning and drive confidence toward near-binary values. Verbalized confidence also provides no explicit assessment of visual support. We therefore separate capability learning from confidence estimation and propose \textbf{DualRead}. DualRead builds on the insight that reliability can be read from the actor's internal states at critical moments in the answering process. It freezes the GRPO-trained actor and combines pre-answer solvability with a post-answer assessment of the generated answer and its visual support. To further assess whether confidence reflects visual grounding, we introduce \textbf{Counterfactual Confidence Grounding AUC} (CCG-AUC). It measures whether confidence decreases when real-image substitution changes the actor from correct to incorrect. Across two VLM backbones and both in- and out-of-distribution medical VQA benchmarks, DualRead improves correctness discrimination and calibration over verbalized confidence while preserving answer accuracy. CCG-AUC reveals whether confidence responds to answer-relevant visual evidence rather than primarily to non-visual cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。