利用不可靠样本提升医学图像分类的压缩效率与精度
Trust the Unreliability: Inward Backward Dynamic Unreliability Driven Coreset Selection for Medical Image Classification
- 从训练中动态评估样本可靠性,聚焦易被遗忘且信心波动大的样本
- 在高压缩率下仍优于现有方法,显著提升决策边界建模能力
- 适合资源受限场景下的医学图像模型训练,尤其关注边界学习
在资源有限的情况下高效处理大规模医学影像数据面临挑战。尽管核心集选择可降低计算成本,但其在医学数据上的效果受限于内在复杂性,如类内差异大、类间相似度高。我们重新审视训练过程,发现神经网络对靠近类别中心的样本有稳定置信度且更易记忆,但过度依赖这些样本会阻碍决策边界建模。因此,我们认为越不可靠的样本越有助于构建边界。为此,提出动态不可靠性驱动的核心集选择(DUCS)策略:1)向内自省:通过分析置信度变化量化样本不确定性;2)向后记忆追踪:通过遗忘频率评估样本留存能力。选择具有显著置信度波动且反复被遗忘的样本,确保其位于决策边界附近,从而帮助模型精炼边界。在多个公开医学数据集上的实验表明,本方法在高压缩率下显著优于现有最先进方法。
原文摘要 · Abstract (English)
Efficiently managing and utilizing large-scale medical imaging datasets with limited resources presents significant challenges. While coreset selection helps reduce computational costs, its effectiveness in medical data remains limited due to inherent complexity, such as large intra-class variation and high inter-class similarity. To address this, we revisit the training process and observe that neural networks consistently produce stable confidence predictions and better remember samples near class centers in training. However, concentrating on these samples may complicate the modeling of decision boundaries. Hence, we argue that the more unreliable samples are, in fact, the more informative in helping build the decision boundary. Based on this, we propose the Dynamic Unreliability-Driven Coreset Selection(DUCS) strategy. Specifically, we introduce an inward-backward unreliability assessment perspective: 1) Inward Self-Awareness: The model introspects its behavior by analyzing the evolution of confidence during training, thereby quantifying uncertainty of each sample. 2) Backward Memory Tracking: The model reflects on its training tracking by tracking the frequency of forgetting samples, thus evaluating its retention ability for each sample. Next, we select unreliable samples that exhibit substantial confidence fluctuations and are repeatedly forgotten during training. This selection process ensures that the chosen samples are near the decision boundary, thereby aiding the model in refining the boundary. Extensive experiments on public medical datasets demonstrate our superior performance compared to state-of-the-art(SOTA) methods, particularly at high compression rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。