arXiv:2409.10544cs.CV2024-09中稿 · IMVIP 2024被引 13

用填充增强与集成学习,解决癌症分类数据少且不均衡问题

OxML Challenge 2023: Carcinoma classification using data augmentation

  • 用五网络集成+图像填充增强应对小样本和不平衡数据
  • 在挑战赛中进入前三,成为冠军团队
  • 适合医疗图像少样本分类任务的研究者参考

癌瘤是主要的癌症类型,可出现在身体多个部位,但医学数据常因隐私限制而稀缺且高度不平衡,正类样本少、负类样本多。OxML 2023 挑战赛提供的数据集规模小且严重失衡,对癌瘤分类构成重大挑战。参赛者普遍采用预训练模型、数据预处理和少样本学习方法。本文提出一种新方法,结合填充数据增强与神经网络集成,利用五个神经网络的集成结构,并针对不同图像尺寸实施填充增强,以提升分类性能。实验表明该方法有效,在挑战赛中位列前三并夺得冠军。

原文摘要 · Abstract (English)

Carcinoma is the prevailing type of cancer and can manifest in various body parts. It is widespread and can potentially develop in numerous locations within the body. In the medical domain, data for carcinoma cancer is often limited or unavailable due to privacy concerns. Moreover, when available, it is highly imbalanced, with a scarcity of positive class samples and an abundance of negative ones. The OXML 2023 challenge provides a small and imbalanced dataset, presenting significant challenges for carcinoma classification. To tackle these issues, participants in the challenge have employed various approaches, relying on pre-trained models, preprocessing techniques, and few-shot learning. Our work proposes a novel technique that combines padding augmentation and ensembling to address the carcinoma classification challenge. In our proposed method, we utilize ensembles of five neural networks and implement padding as a data augmentation technique, taking into account varying image sizes to enhance the classifier's performance. Using our approach, we made place into top three and declared as winner.

癌症分类数据增强少样本学习医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。