提升胸部X光片罕见病多标签零样本分类准确率
CXR-CML: Improved zero-shot classification of long-tailed multi-label diseases in Chest X-Rays
- 用高斯混合模型聚类特征空间,动态调整类别权重
- 在MIMIC-CXR-JPG数据集上零样本AUC提升7个百分点
- 特别改善罕见病识别,适合医疗影像诊断研究者
胸部X光检查在疾病诊断中至关重要,但临床表现的类别分布存在严重不平衡,导致当前自监督深度学习模型难以准确识别长尾类别。尽管对比语言图像预训练(CLIP)等视觉-语言模型能有效建模潜在空间流形,实现高零样本分类精度,但在长尾类别上性能显著下降。本文提出一种基于潜在空间类别分布的加权机制,通过高斯混合模型(GMM)对特征进行聚类,并利用学生t分布进一步优化聚类结果,结合度量损失调整嵌入表示。该方法实现特征的稳定自适应聚类,在MIMIC-CXR-JPG数据集40个类别上相较先前最优模型平均提升7%零样本AUC,显著增强罕见病识别能力。
原文摘要 · Abstract (English)
Chest radiography (CXR) plays a crucial role in the diagnosis of various diseases. However, the inherent class imbalance in the distribution of clinical findings presents a significant challenge for current self-supervised deep learning models. These models often fail to accurately classify long-tailed classes. Current Vision-Language models such as Contrastive Language Image Pre-training (CLIP) models effectively model the manifold distribution of the latent space, enabling high zero-shot classification accuracies. Although CLIP performs well on most of the primary classes in the dataset, our work reveals that its effectiveness decreases significantly for classes with a long-tailed distribution. Our approach employs a class-weighting mechanism that directly aligns with the distribution of classes within the latent space. This method ensures a substantial improvement in overall classification performance, with particular emphasis on enhancing the recognition and accuracy of rarely observed classes. We accomplish this by applying Gaussian Mixture Model (GMM) clustering to the latent space. The subsequent clusters are further refined by Student t-distribution, followed by a metric loss that utilizes the altered embeddings. Our approach facilitates stable and adaptive clustering of the features. This results in a notable average improvement of 7\% points in zero-shot AUC scores across 40 classes in the MIMIC-CXR-JPG dataset from previous SOTA models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。