arXiv:2512.22872cs.CV2025-12

Lamps通过多视角自监督学习,让模型更懂人体解剖结构。

Lamps: Learning Anatomy from Multiple Perspectives via Self-supervision in Chest Radiographs

  • 利用解剖一致性、连贯性和层级性作为自监督信号
  • 在10个数据集上表现优于10个基线模型
  • 适合医学影像基础模型研究者参考

基础模型在自然语言处理和计算机视觉中取得成功,因其能捕捉自然语言的底层结构。但在医学影像中,核心基础是人体解剖结构,因为影像直接反映身体内部构造,体现解剖的一致性、连贯性和层级性。然而,现有自监督学习方法常忽视这些特性,限制了对解剖特征的有效学习。为此,我们构建了Lamps(通过多视角自监督学习解剖结构),在大规模胸片数据上预训练,和谐利用解剖的一致性、连贯性和层级性作为监督信号。在10个数据集上的广泛实验,通过微调和涌现属性分析表明,相比10个基线模型,Lamps展现出更优的鲁棒性、可迁移性和临床潜力。通过多视角学习,Lamps为基础模型发展与人体解剖结构一致的有意义、稳健表示提供了独特机会。

原文摘要 · Abstract (English)

Foundation models have been successful in natural language processing and computer vision because they are capable of capturing the underlying structures (foundation) of natural languages. However, in medical imaging, the key foundation lies in human anatomy, as these images directly represent the internal structures of the body, reflecting the consistency, coherence, and hierarchy of human anatomy. Yet, existing self-supervised learning (SSL) methods often overlook these perspectives, limiting their ability to effectively learn anatomical features. To overcome the limitation, we built Lamps (learning anatomy from multiple perspectives via self-supervision) pre-trained on large-scale chest radiographs by harmoniously utilizing the consistency, coherence, and hierarchy of human anatomy as the supervision signal. Extensive experiments across 10 datasets evaluated through fine-tuning and emergent property analysis demonstrate Lamps' superior robustness, transferability, and clinical potential when compared to 10 baseline models. By learning from multiple perspectives, Lamps presents a unique opportunity for foundation models to develop meaningful, robust representations that are aligned with the structure of human anatomy.

医学影像自监督学习解剖结构基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。