构建首个统一医疗图像长尾分类基准,支持30+方法在12个数据集上对比。
MONICA: Benchmarking on Long-tailed Medical Image Classification
- 整合30余种长尾学习方法,覆盖6大医学领域
- 在12个真实医疗数据集上系统评估性能差异
- 开源代码库助力可复现研究,适合医疗AI开发者
长尾学习在数据不平衡场景中极具挑战性,旨在从遵循长尾分布的大量医学影像中训练出泛化能力强的模型。皮肤镜检查和胸部X光等诊断影像常呈现复杂的长尾临床发现分布。尽管近年来医学图像分析中的长尾学习受到广泛关注,但该领域仍缺乏统一、严谨且全面的基准测试体系,导致比较不公、结论模糊。为此,我们构建了名为医学开源长尾分类(MONICA)的统一、结构化的代码库,集成了超过30种相关领域的先进方法,并在涵盖6个医学领域的12个长尾医疗数据集上进行了评估。本工作为该领域提供了实用指导与深入分析,揭示了主流方法中各组件的有效性。我们期望该代码库能成为可复现的综合性基准,推动长尾医疗图像学习的进一步发展。代码已公开于 https://github.com/PyJulie/MONICA。
原文摘要 · Abstract (English)
Long-tailed learning is considered to be an extremely challenging problem in data imbalance learning. It aims to train well-generalized models from a large number of images that follow a long-tailed class distribution. In the medical field, many diagnostic imaging exams such as dermoscopy and chest radiography yield a long-tailed distribution of complex clinical findings. Recently, long-tailed learning in medical image analysis has garnered significant attention. However, the field currently lacks a unified, strictly formulated, and comprehensive benchmark, which often leads to unfair comparisons and inconclusive results. To help the community improve the evaluation and advance, we build a unified, well-structured codebase called Medical OpeN-source Long-taIled ClassifiCAtion (MONICA), which implements over 30 methods developed in relevant fields and evaluated on 12 long-tailed medical datasets covering 6 medical domains. Our work provides valuable practical guidance and insights for the field, offering detailed analysis and discussion on the effectiveness of individual components within the inbuilt state-of-the-art methodologies. We hope this codebase serves as a comprehensive and reproducible benchmark, encouraging further advancements in long-tailed medical image learning. The codebase is publicly available on https://github.com/PyJulie/MONICA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。