用扩散模型生成病变热图,提升医学图像检索与分类性能
DAug: Diffusion-based Channel Augmentation for Radiology Image Retrieval and Classification
- 通过扩散模型生成疾病易发区域的热图,扩充图像通道
- 在多个数据集上实现最优的检索与分类准确率
- 适合需要提升小样本医学图像理解能力的研究者
医学图像理解需细致关注细微视觉特征,特定区域更需重点关注。尽管放射科医生经过多年经验积累形成专业判断力,但人工智能模型在训练数据有限的情况下难以学习何处应重点观察,导致医学图像理解的鲁棒性不足。为此,我们提出基于扩散的特征增强方法(DAug),一种可移植的方案,利用生成模型输出提升感知模型性能。具体而言,将放射科图像扩展为多通道图像,新增通道为疾病易发区域的热图,采用条件扩散图像到图像转换模型生成这些热图,输入为选定疾病类别。该方法基于生成模型学习正常与异常图像分布的特性,其知识可补充图像理解任务。此外,提出图像-文本-类别混合对比学习,同时利用文本和类别标签。结合两项新方法,无需改变模型结构即可超越基线模型,在医学图像检索与分类任务中达到当前最优表现。
原文摘要 · Abstract (English)
Medical image understanding requires meticulous examination of fine visual details, with particular regions requiring additional attention. While radiologists build such expertise over years of experience, it is challenging for AI models to learn where to look with limited amounts of training data. This limitation results in unsatisfying robustness in medical image understanding. To address this issue, we propose Diffusion-based Feature Augmentation (DAug), a portable method that improves a perception model's performance with a generative model's output. Specifically, we extend a radiology image to multiple channels, with the additional channels being the heatmaps of regions where diseases tend to develop. A diffusion-based image-to-image translation model was used to generate such heatmaps conditioned on selected disease classes. Our method is motivated by the fact that generative models learn the distribution of normal and abnormal images, and such knowledge is complementary to image understanding tasks. In addition, we propose the Image-Text-Class Hybrid Contrastive learning to utilize both text and class labels. With two novel approaches combined, our method surpasses baseline models without changing the model architecture, and achieves state-of-the-art performance on both medical image retrieval and classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。