仅用临床报告文本训练模型,实现小样本下心脏病灶检测
Fake It Till You Make It: Using Synthetic Data and Domain Knowledge for Improved Text-Based Learning for LGE Detection
- 基于医学知识生成合成病灶图像与对应文本,扩充小样本数据
- 通过解剖导向图像定向与字幕损失,提升图文对齐精度
- 适用于标注稀缺但报告丰富的医疗影像分析场景
心脏延迟增强磁共振(LGE MRI)中瘢痕的高强化区域检测是一项需要大量临床经验的复杂任务。尽管深度学习模型已展现出良好效果,但其训练需大量带精细标注的数据。临床报告包含病变位置、范围及病因等丰富信息。虽然基于CLIP的预训练可利用图文对齐,但仍需大规模数据和下游微调。本研究在仅965例患者的有限队列上,仅使用临床报告文本,结合领域知识设计方法:通过系统性生成病灶图像与对应文本实现数据增强;采用解剖学指导的图像方向标准化以改善空间与文本特征对齐;引入字幕损失实现细粒度监督,并探索视觉编码器预训练的影响。最终通过消融实验验证各组件贡献。结果表明,该方法可在小样本下有效提升检测性能。
原文摘要 · Abstract (English)
Detection of hyperenhancement from cardiac LGE MRI images is a complex task requiring significant clinical expertise. Although deep learning-based models have shown promising results for the task, they require large amounts of data with fine-grained annotations. Clinical reports generated for cardiac MR studies contain rich, clinically relevant information, including the location, extent and etiology of any scars present. Although recently developed CLIP-based training enables pretraining models with image-text pairs, it requires large amounts of data and further finetuning strategies on downstream tasks. In this study, we use various strategies rooted in domain knowledge to train a model for LGE detection solely using text from clinical reports, on a relatively small clinical cohort of 965 patients. We improve performance through the use of synthetic data augmentation, by systematically creating scar images and associated text. In addition, we standardize the orientation of the images in an anatomy-informed way to enable better alignment of spatial and text features. We also use a captioning loss to enable fine-grained supervision and explore the effect of pretraining of the vision encoder on performance. Finally, ablation studies are carried out to elucidate the contributions of each design component to the overall performance of the model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。