arXiv:2607.13800eess.IVcs.CV2026-07

研究胸部X光片多图分类中临床指征与报告的泄漏问题,提出更可靠的融合方法。

Prospective clinical indication, post-hoc report leakage, and fusion design in multi-image chest radiograph classification: a patient-clustered evaluation

论文配图:Prospective clinical indication, post-hoc report leakage, and fusion design in multi-image chest radiograph classification: a patient-clustered evaluation
图 1 · 摘自论文原文
  • 设计了感知图像顺序的融合模型,避免报告信息泄露干扰
  • 发现临床指征可显著提升分类性能,但需防事后报告污染
  • 适合医疗AI研究者、临床数据建模人员参考

胸部X光数据集常包含多个图像及临床指征、发现和印象等信息,但这些输入产生于诊疗不同阶段。本研究基于15,000例ReXGradient-160K病例,每例含两张可读图像与五项CheXbert衍生报告观察。使用Frozen DenseNet-121与Bio+ClinicalBERT编码器,对比图像仅用、指征仅用、固定顺序多模态、随机交换、DeepSets及SectionGuard-MI模型。发现与印象仅作为事后泄漏控制。模型训练采用五组随机种子,通过2,000次患者聚类自助重复估计公开测试不确定性。在U-Ones下,单图宏观AUROC为0.643,双图达0.694,指征达0.749,普通双图+指征融合达0.780。SectionGuard-MI获AUROC 0.783、AUPRC 0.260;相较于普通融合,其配对AUROC差异为0.0031(95% CI: -0.0042~0.0104;校正p=0.374),AUPRC差异为0.0289(95% CI: 0.0095~0.0413;校正p=0.004)。DeepSets在前瞻性评估中取得最高AUROC点估计(0.787),随机交换融合的前瞻性AUPRC点估计最高(0.265),且校准性优于SectionGuard-MI。仅使用完整报告文本时,AUROC达0.979,AUPRC达0.836;即使精确或扩展掩码后,AUROC仍高于0.973。结果表明:前瞻性临床指征与报告目标强相关,排列感知融合具竞争力,而事后报告文本引发严重报告-标签循环。

原文摘要 · Abstract (English)

Chest radiograph datasets often combine multiple images with Clinical Indication, Findings, and Impression, although these inputs are produced at different stages of care. We evaluated 15,000 ReXGradient-160K studies with two readable images and five CheXbert-derived report observations. Frozen DenseNet-121 and Bio+ClinicalBERT encoders were used to compare image-only, Indication-only, fixed-order multimodal, random-swap, DeepSets, and SectionGuard-MI models. Findings and Impression were evaluated only as post-hoc leakage controls. Models were trained with five seeds, and public-test uncertainty was estimated with 2,000 patient-cluster bootstrap replicates. Under U-Ones, macro AUROC was 0.643 for the primary image, 0.694 for two images, 0.749 for Indication, and 0.780 for ordinary two-image-plus-Indication fusion. SectionGuard-MI achieved AUROC 0.783 and AUPRC 0.260. Relative to ordinary fusion, its paired AUROC difference was 0.0031 (95% CI, -0.0042 to 0.0104; adjusted p=0.374), while its AUPRC difference was 0.0289 (95% CI, 0.0095 to 0.0413; adjusted p=0.004). DeepSets had the highest prospective AUROC point estimate (0.787), and random-swap fusion had the highest prospective AUPRC point estimate (0.265) with better calibration than SectionGuard-MI. Full report text alone reached AUROC 0.979 and AUPRC 0.836; AUROC remained above 0.973 after exact or expanded masking. These results show that prospective Indication is strongly associated with report-derived targets, permutation-aware fusion is competitive, and post-hoc report text creates substantial report-label circularity.

医学影像多模态融合报告泄漏临床指征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。