用多尺度引导的生成预训练,提升医学图像关键点检测在少量标注下的准确与可靠
CDPM-Align: Multi-Scale Guidance-Aligned Diffusion Pretraining for Robust Few-Shot Anatomical Landmark Detection

- 通过条件扩散模型在小样本上进行多尺度生成预训练
- 仅用10或25张标注图像,仍实现亚毫米级定位精度与低不确定性
- 适合资源受限但需高可靠性的临床医学图像分析场景
解剖学关键点检测是医学影像分析的基础任务,支持广泛的诊断与介入流程。尽管近期方法已实现亚毫米级定位精度,但临床部署还需预测的可靠性与鲁棒性。尽管其临床意义重大,该领域的表征学习影响仍待深入探索。本文提出CDPM-Align,一种用于解剖学关键点检测的多尺度引导对齐条件扩散预训练方法。实验基于少量图像和少量标注设置,采用三个流行的异构小规模基准数据集进行条件生成式预训练。下游任务中考虑低标注场景,使用10和25张标注图像,反映临床标注工作量与资源约束间的现实权衡。结果表明,生成式预训练使模型学习到稳健表征,显著提升下游任务的准确率与不确定性估计性能,推动安全高效的临床应用。
原文摘要 · Abstract (English)
Anatomical landmark detection is a fundamental task in medical image analysis supporting a wide range of diagnostic and interventional workflows. Although recent methods have achieved sub-millimetric localisation, accuracy alone is not sufficient for clinical deployment, requiring reliability and robustness in prediction. Despite its clinical relevance, the impact of representation learning in this context is still underexplored. In this work, we introduce CDPM-align, a multi-scale guidance-aligned conditional diffusion pre-training for anatomical landmark detection. Our experimental setup focuses on a few images and a few annotation regimes. Specifically, we employ three popular heterogeneous small-scale benchmark datasets for representation learning via conditional generative pre-training. Furthermore, we consider low-annotation scenarios for the downstream task of landmark detection, with 10 and 25 annotated images, reflecting realistic trade-offs between clinical effort and resource constraints for annotations. Our results confirm that generative pre-training enables the model to learn a robust representation. This improves both accuracy and uncertainty on the downstream tasks, advancing towards safe and efficient clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。