用渐进式深度学习提升颅底软骨融合阶段的自动评估精度
Progressive Deep Learning for Automated Spheno-Occipital Synchondrosis Maturation Assessment
- 模仿临床医生从宏观到细微的判断思路,分步训练网络
- 在中间过渡阶段准确率显著提升,优化更稳定
- 无需改架构或损失函数,适合医学影像连续过程建模
精准评估蝶枕软骨(SOS)成熟度是判断颅面生长和正畸/手术时机的关键。然而,基于锥形束CT(CBCT)的SOS分期依赖于细微且持续演变的形态学特征,导致观察者间差异大、可重复性差,尤其在融合过渡阶段。本文将SOS评估视为细粒度视觉识别问题,提出一种渐进式表征学习框架,模拟专家临床推理:从粗略颅底结构逐步聚焦到闭合的细微模式。不采用端到端全容量训练,而是按时间顺序逐层激活深层模块,使浅层先编码稳定的颅底形态,深层再专注于区分相邻成熟阶段。该策略形成随网络深度演进的课程学习,与SOS融合的生物连续性对齐。跨卷积与Transformer架构的实验表明,此方法相比标准训练获得更稳定优化和更高精度,尤其在模糊中间阶段表现突出。重要的是,仅通过调整训练动态即实现性能提升,无需修改网络结构或损失函数,证明训练机制本身即可显著改善解剖表征学习。该框架建立了临床直觉与深度视觉表示间的原理性联系,支持从CBCT中鲁棒、数据高效地进行SOS分期,为医学影像中其他连续生物过程建模提供通用策略。
原文摘要 · Abstract (English)
Accurate assessment of spheno-occipital synchondrosis (SOS) maturation is a key indicator of craniofacial growth and a critical determinant for orthodontic and surgical timing. However, SOS staging from cone-beam CT (CBCT) relies on subtle, continuously evolving morphological cues, leading to high inter-observer variability and poor reproducibility, especially at transitional fusion stages. We frame SOS assessment as a fine-grained visual recognition problem and propose a progressive representation-learning framework that explicitly mirrors how expert clinicians reason about synchondral fusion: from coarse anatomical structure to increasingly subtle patterns of closure. Rather than training a full-capacity network end-to-end, we sequentially grow the model by activating deeper blocks over time, allowing early layers to first encode stable cranial base morphology before higher-level layers specialize in discriminating adjacent maturation stages. This yields a curriculum over network depth that aligns deep feature learning with the biological continuum of SOS fusion. Extensive experiments across convolutional and transformer-based architectures show that this expert-inspired training strategy produces more stable optimization and consistently higher accuracy than standard training, particularly for ambiguous intermediate stages. Importantly, these gains are achieved without changing network architectures or loss functions, demonstrating that training dynamics alone can substantially improve anatomical representation learning. The proposed framework establishes a principled link between expert dental intuition and deep visual representations, enabling robust, data-efficient SOS staging from CBCT and offering a general strategy for modeling other continuous biological processes in medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。