arXiv:2603.29492cs.CL2026-03

让AI生成的放射科报告自带可信度标注,避免误判影响诊疗。

Calibrated Confidence Expression for Radiology Report Generation

  • 用强化学习训练模型,生成带可信度的报告,分整体和句子级评分。
  • 实验显示可信度与医生判断高度一致,显著优于现有方法。
  • 适合需安全部署AI辅助报告生成的医院和临床场景。

大型视觉-语言模型(LVLM)在放射科报告生成中的安全应用,不仅需要准确预测,还需提供可解释的置信度指标,以指导医生重点审查,降低幻觉风险。现有模型常过度自信,且多模态环境下置信度校准研究匮乏。为此,我们提出ConRad(放射科报告置信度校准)框架,通过强化学习微调医学LVLM,生成校准后的口头置信度评估。研究两种设置:单报告级置信度和逐句级置信度,均采用GRPO算法,基于对数评分规则设计奖励函数,鼓励真实自评并保证最优校准。实验表明,ConRad显著提升置信度校准效果,优于对比方法。临床评估显示,其报告级评分与医生判断高度一致。通过标记整份报告或低置信度语句,支持更安全的AI辅助报告生成临床集成。

原文摘要 · Abstract (English)

Safe deployment of Large Vision-Language Models (LVLMs) in radiology report generation requires not only accurate predictions but also clinically interpretable indicators of when outputs should be thoroughly reviewed, enabling selective radiologist verification and reducing the risk of hallucinated findings influencing clinical decisions. One intuitive approach to this is verbalized confidence, where the model explicitly states its certainty. However, current state-of-the-art language models are often overconfident, and research on calibration in multimodal settings such as radiology report generation is limited. To address this gap, we introduce ConRad (Confidence Calibration for Radiology Reports), a reinforcement learning framework for fine-tuning medical LVLMs to produce calibrated verbalized confidence estimates alongside radiology reports. We study two settings: a single report-level confidence score and a sentence-level variant assigning a confidence to each claim. Both are trained using the GRPO algorithm with reward functions based on the logarithmic scoring rule, which incentivizes truthful self-assessment by penalizing miscalibration and guarantees optimal calibration under reward maximization. Experimentally, ConRad substantially improves calibration and outperforms competing methods. In a clinical evaluation we show that ConRad's report level scores are well aligned with clinicians' judgment. By highlighting full reports or low-confidence statements for targeted review, ConRad can support safer clinical integration of AI-assistance for report generation.

放射科报告置信度校准医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。