给医学影像报告生成加不确定度评估,提升临床可信度。
CONRep: Uncertainty-Aware Vision-Language Report Drafting Using Conformal Prediction
- 用置信区间方法对影像报告的结论和语句分级,量化模型不确定性。
- 高置信度输出与放射科医生标注一致率显著高于低置信度结果。
- 无需修改原有模型,可适配各类视觉语言模型,适合临床部署。
使用视觉语言模型(VLMs)进行自动化放射科报告生成(ARRD)进展迅速,但多数系统缺乏显式的不确定度估计,限制了其可信度和安全临床应用。本文提出CONRep,一种与模型无关的框架,通过结合置信区间(CP)为VLM生成的放射科报告提供统计上可靠的不确定度量化。CONRep在标签层面校准预定义病灶的二分类预测,在句子层面通过图像-文本语义对齐评估自由文本描述的不确定性。我们在公开的胸部X光数据集上,使用生成式和对比式VLM进行了评估。在两种设置下,高置信度输出与放射科医生标注及真实报告的一致性均显著高于低置信度输出。CONRep通过无需修改底层模型即可实现校准的置信度分层,提升了自动化放射科报告系统的透明性、可靠性与临床可用性。
原文摘要 · Abstract (English)
Automated radiology report drafting (ARRD) using vision-language models (VLMs) has advanced rapidly, yet most systems lack explicit uncertainty estimates, limiting trust and safe clinical deployment. We propose CONRep, a model-agnostic framework that integrates conformal prediction (CP) to provide statistically grounded uncertainty quantification for VLM-generated radiology reports. CONRep operates at both the label level, by calibrating binary predictions for predefined findings, and the sentence level, by assessing uncertainty in free-text impressions via image-text semantic alignment. We evaluate CONRep using both generative and contrastive VLMs on public chest X-ray datasets. Across both settings, outputs classified as high confidence consistently show significantly higher agreement with radiologist annotations and ground-truth impressions than low-confidence outputs. By enabling calibrated confidence stratification without modifying underlying models, CONRep improves the transparency, reliability, and clinical usability of automated radiology reporting systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。