用医院影像和报告训练自监督模型,诊断不准但预测病情进展更准。
Self-Supervised Learning for Knee Osteoarthritis: Diagnostic Limitations and Prognostic Value of Hospital Data
- 用医院影像和放射科描述做多模态自监督预训练
- 在4年结构进展预测上超越ImageNet基线(外部验证AUC 0.701)
- 适合关注疾病进展预测的研究者,不适用于诊断分级
本研究评估自监督学习(SSL)是否优于ImageNet预训练,在膝骨关节炎(OA)的诊断与预后建模中表现更好。比较了两种方式:(i)仅图像的SSL,基于OAI、MOST和NYU队列的膝关节X光片预训练;(ii)多模态图像-文本SSL,基于医院膝关节X光片及其放射科报告配对数据预训练。在诊断性Kellgren-Lawrence(KL)分级预测上,SSL结果不一致:图像仅预训练在冻结编码器时提升准确率,但在全量微调下未超越ImageNet。多模态SSL同样未改善分级性能。原因可能是预训练数据集中仅包含临床确诊的OA患者,缺乏从正常到重度的完整疾病谱,且放射科报告未明确标注KL等级,难以提供精准监督信号。相反,该多模态初始化显著提升了预后建模效果,在4年结构性病变发生与进展预测中优于ImageNet基线,包括在外部验证集上(MOST队列,10%标签数据下AUROC达0.701,对比基线0.599)。结果表明,尽管医院数据不适于诊断分级,但若下游任务与数据分布匹配,则可为预后建模提供强信号。
原文摘要 · Abstract (English)
This study assesses whether self-supervised learning (SSL) improves knee osteoarthritis (OA) modeling for diagnosis and prognosis relative to ImageNet-pretrained initialization. We compared (i) image-only SSL pretrained on knee radiographs from the OAI, MOST, and NYU cohorts, and (ii) multimodal image-text SSL pretrained on hospital knee radiographs paired with radiologist impressions. For diagnostic Kellgren-Lawrence (KL) grade prediction, SSL yielded mixed results. While image-only SSL improved accuracy during linear probing (frozen encoder), it did not outperform ImageNet pretraining during full fine-tuning. Similarly, multimodal SSL failed to improve grading performance. A likely explanation is mismatch between the hospital pretraining corpus and the downstream diagnostic task: the hospital image-text dataset was restricted to knees from patients with clinically identified OA in routine care, rather than a cohort spanning the full spectrum from normal to severe disease needed for balanced KL grading. In addition, radiology impressions do not explicitly encode KL grade, limiting supervision for learning KL-specific decision boundaries. In contrast, this same multimodal initialization significantly improved prognostic modeling. It outperformed ImageNet baselines in predicting 4-year structural incidence and progression, including on external validation (MOST AUROC: 0.701 vs. 0.599 at 10\% labeled data). Overall, these results suggest that our hospital image-text data may be less effective for diagnostic grading when the pretraining cohort is limited to OA knees, but can provide a strong signal for prognostic modeling when the downstream task is better aligned with the pretraining data distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。