用解剖结构预训练提升医学报告定位精度
Anatomical grounding pre-training for medical phrase grounding
- 以解剖术语与图像区域对齐作为预训练任务
- 在MS-CXR上达到61.2的mIoU,优于现有模型
- 适合医学影像与报告对齐研究者参考
医学短语定位(MPG)旨在将医学报告中描述的放射学发现映射到医学图像中的特定区域。当前进展受限于标注数据稀缺。本文提出解剖结构对齐作为域内预训练任务,利用大规模数据集Chest ImaGenome,将解剖术语与对应图像区域对齐。在MS-CXR上的实证评估表明,该预训练策略在零样本学习和微调设置下均显著提升性能,优于现有最优的MPG模型。微调后模型在MS-CXR上取得61.2的mIoU,验证了解剖结构预训练的有效性。
原文摘要 · Abstract (English)
Medical Phrase Grounding (MPG) maps radiological findings described in medical reports to specific regions in medical images. The primary obstacle hindering progress in MPG is the scarcity of annotated data available for training and validation. We propose anatomical grounding as an in-domain pre-training task that aligns anatomical terms with corresponding regions in medical images, leveraging large-scale datasets such as Chest ImaGenome. Our empirical evaluation on MS-CXR demonstrates that anatomical grounding pre-training significantly improves performance in both a zero-shot learning and fine-tuning setting, outperforming state-of-the-art MPG models. Our fine-tuned model achieved state-of-the-art performance on MS-CXR with an mIoU of 61.2, demonstrating the effectiveness of anatomical grounding pre-training for MPG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。