用病变表现指导负样本挖掘,提升乳腺钼靶分类效果。
ManiNeg: Manifestation-guided Multimodal Pretraining for Mammography Classification
- 以病变表现为先验,自动筛选难负样本,增强特征表达
- 在多模态和单模态下均提升分类性能,跨数据集泛化良好
- 开源了带标注病变的多视角钼靶数据集,助力后续研究
乳腺癌是威胁人类健康的重大疾病。对比学习已成为从钼靶图像中提取关键病灶特征的有效方法,为乳腺癌筛查与分析提供有力支持。对比学习中的关键环节是负样本采样,合理选择困难负样本对保留病灶细节信息至关重要。传统方法通常假设特征可充分捕捉语义内容,且每个小批量样本天然包含理想难负样本,但乳腺肿块特性挑战了这一假设。为此,我们提出ManiNeg,利用病变表现(即疾病的可见症状或体征)作为代理,挖掘难负样本。病变表现提供基于知识且鲁棒的负样本选择依据,且不受模型优化影响,采样效率高。为支持ManiNeg及未来研究,我们构建了MVKL数据集,包含多视角钼靶图像、对应报告、精心标注的病变表现及病理确诊的良恶性结果。我们在良恶性分类任务上评估ManiNeg,结果表明该方法不仅提升了单模态与多模态下的表征能力,还展现出跨数据集的良好泛化性。代码与数据集已公开于https://github.com/wxwxwwxxx/ManiNeg。
原文摘要 · Abstract (English)
Breast cancer is a significant threat to human health. Contrastive learning has emerged as an effective method to extract critical lesion features from mammograms, thereby offering a potent tool for breast cancer screening and analysis. A crucial aspect of contrastive learning involves negative sampling, where the selection of appropriate hard negative samples is essential for driving representations to retain detailed information about lesions. In contrastive learning, it is often assumed that features can sufficiently capture semantic content, and that each minibatch inherently includes ideal hard negative samples. However, the characteristics of breast lumps challenge these assumptions. In response, we introduce ManiNeg, a novel approach that leverages manifestations as proxies to mine hard negative samples. Manifestations, which refer to the observable symptoms or signs of a disease, provide a knowledge-driven and robust basis for choosing hard negative samples. This approach benefits from its invariance to model optimization, facilitating efficient sampling. To support ManiNeg and future research endeavors, we developed the MVKL dataset, which includes multi-view mammograms, corresponding reports, meticulously annotated manifestations, and pathologically confirmed benign-malignant outcomes. We evaluate ManiNeg on the benign and malignant classification task. Our results demonstrate that ManiNeg not only improves representation in both unimodal and multimodal contexts but also shows generalization across datasets. The MVKL dataset and our codes are publicly available at https://github.com/wxwxwwxxx/ManiNeg.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。