在小样本医学图像中,深度模型分割效果不稳定,排名不可靠。
Challenges in Deep Learning-Based Small Organ Segmentation: A Benchmarking Perspective for Medical Research with Limited Datasets
- 对比10种模型在9张病理图上表现,发现经典模型易失效
- 跨数据集测试时模型排名大幅变化,性能不一致
- 建议用不确定性评估替代传统基准,尤其适用于小样本研究
准确分割颈动脉组织结构对心血管疾病研究至关重要。本研究在仅9张心血管病理图像的有限数据集上,系统评估了包括经典架构、现代CNN、视觉Transformer及基础模型在内的十种深度学习分割模型。通过消融实验分析数据增强、输入分辨率和随机种子稳定性对结果的影响。在独立泛化数据集(N=153)上进行分布外测试发现,基础模型保持性能而经典架构失效,且模型排名在分布内与分布外场景间差异显著。在不同样本量下训练同一模型,揭示了数据集特有的排名层级,表明模型排名不具备跨数据集通用性。尽管采用贝叶斯超参数优化,模型性能仍高度依赖数据划分。自助法分析显示顶级模型置信区间重叠严重,差异主要由统计噪声而非算法优势驱动。该不稳定性暴露了低数据临床场景下标准基准的局限性,挑战了性能排名反映临床实用性的假设。我们从两方面倡导不确定性感知评估:一是此类场景广泛存在;二是有助于在早期即判断研究方向是否值得继续。
原文摘要 · Abstract (English)
Accurate segmentation of carotid artery structures in histopathological images is vital for cardiovascular disease research. This study systematically evaluates ten deep learning segmentation models including classical architectures, modern CNNs, a Vision Transformer, and foundation models, on a limited dataset of nine cardiovascular histology images. We conducted ablation studies on data augmentation, input resolution, and random seed stability to quantify sources of variance. Evaluation on an independent generalization dataset ($N=153$) under distribution shift reveals that foundation models maintain performance while classical architectures fail, and that rankings change substantially between in-distribution and out-of-distribution settings. Training on the second dataset at varying sample sizes reveals dataset-specific ranking hierarchies confirming that model rankings are not generalizable across datasets. Despite rigorous Bayesian hyperparameter optimization, model performance remains highly sensitive to data splits. The bootstrap analysis reveals substantially overlapping confidence intervals among top models, with differences driven more by statistical noise than algorithmic superiority. This instability exposes limitations of standard benchmarking in low-data clinical settings and challenges assumptions that performance rankings reflect clinical utility. We advocate for uncertainty-aware evaluation in low-data clinical research scenarios from two perspectives. First, the scenario is not niche and is rather widely spread; and second, it enables pursuing or discontinuing research tracks with limited datasets from incipient stages of observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。