arXiv:2601.18250cs.CVcs.AI2026-01

基于120万张影像训练的多模态模型,可通用诊断膝髋肩踝病灶。

A multimodal vision foundation model for generalizable knee pathology

  • 用自监督对比学习在120万张膝关节X光与MRI上预训练
  • 在14项任务中达顶尖性能,仅需50%标注数据即超监督模型
  • 跨关节通用性强,膝部训练后可准确识别髋肩踝病变

肌肉骨骼疾病是全球致残的主要原因,亟需精准的医学影像解读。当前骨科人工智能多依赖特定任务、有监督学习,存在数据碎片化、标注需求大、泛化能力差等问题。由于缺乏大规模、高质量、开源的肌肉骨骼数据集,基础模型发展受限。为此,我们提出OrthoFoundation,一个面向肌肉骨骼病理的多模态视觉基础模型。构建了包含120万张未标注膝关节X光与MRI图像的预训练数据集,采用Dinov3骨干网络,通过自监督对比学习捕捉鲁棒的放射学表征。OrthoFoundation在14项下游任务中达到领先水平,膝关节X光骨关节炎诊断准确率优异,磁共振结构损伤检测排名第一。模型表现出显著标签效率,仅用50%标注数据即超越传统监督基线。尽管预训练仅针对膝部影像,其跨解剖结构泛化能力突出,可有效应用于髋、肩、踝等部位。该模型为通用性肌肉骨骼影像AI提供新范式,通过大规模多模态数据学习共通的放射学语义,克服传统模型局限,降低标注负担,提升临床诊断准确性。

原文摘要 · Abstract (English)

Musculoskeletal disorders represent a leading cause of global disability, creating an urgent demand for precise interpretation of medical imaging. Current artificial intelligence (AI) approaches in orthopedics predominantly rely on task-specific, supervised learning paradigms. These methods are inherently fragmented, require extensive annotated datasets, and often lack generalizability across different modalities and clinical scenarios. The development of foundation models in this field has been constrained by the scarcity of large-scale, curated, and open-source musculoskeletal datasets. To address these challenges, we introduce OrthoFoundation, a multimodal vision foundation model optimized for musculoskeletal pathology. We constructed a pre-training dataset of 1.2 million unlabeled knee X-ray and MRI images from internal and public databases. Utilizing a Dinov3 backbone, the model was trained via self-supervised contrastive learning to capture robust radiological representations. OrthoFoundation achieves state-of-the-art (SOTA) performance across 14 downstream tasks. It attained superior accuracy in X-ray osteoarthritis diagnosis and ranked first in MRI structural injury detection. The model demonstrated remarkable label efficiency, matching supervised baselines using only 50% of labeled data. Furthermore, despite being pre-trained on knee images, OrthoFoundation exhibited exceptional cross-anatomy generalization to the hip, shoulder, and ankle. OrthoFoundation represents a significant advancement toward general-purpose AI for musculoskeletal imaging. By learning fundamental, joint-agnostic radiological semantics from large-scale multimodal data, it overcomes the limitations of conventional models, which provides a robust framework for reducing annotation burdens and enhancing diagnostic accuracy in clinical practice.

医学影像基础模型多模态跨关节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。