用视觉语言模型提升膝骨关节炎分级准确率,尤其改善早期阶段判读一致性。
VL-OrdinalFormer: Vision Language Guided Ordinal Transformers for Interpretable Knee Osteoarthritis Grading
- 融合视觉语言先验的序数回归框架,结合文本语义增强图像理解。
- 在OAI数据集上达到最佳宏观F1和准确率,对KL1/2阶段提升显著。
- 通过注意力图验证可解释性,适合临床辅助诊断场景使用。
膝骨关节炎(KOA)是全球主要致残原因,基于凯尔格伦-劳斯(KL)分级系统进行严重程度评估对临床决策至关重要。然而,早期疾病阶段(尤其是KL1与KL2之间)的放射学差异细微,常导致放射科医生间判断不一致。为此,本文提出VLOrdinalFormer,一种用于从膝关节X光片自动分级的视觉语言引导序数学习框架。该方法结合ViT-L16主干网络、基于CORAL的序数回归以及受CLIP驱动的语义对齐模块,使模型能融入关节间隙狭窄、骨赘形成、软骨下硬化等临床相关文本概念。为提升鲁棒性并缓解过拟合,采用分层五折交叉验证、类别感知重加权以强化中间难度等级,并引入测试时增强与全局阈值优化。在公开的OAI kneeKL224数据集上的实验表明,VLOrdinalFormer在宏平均F1和整体准确率上均优于CNN与ViT基线模型。尤其在KL1和KL2阶段表现显著提升,同时未牺牲轻度或重度病例的分类精度。通过Grad CAM与CLIP相似度图的可解释性分析显示,模型始终关注临床相关解剖区域。结果表明,视觉语言对齐的序数变压器具有作为可靠且可解释的KOA分级工具的潜力,适用于常规放射学实践中的疾病进展评估。
原文摘要 · Abstract (English)
Knee osteoarthritis (KOA) is a leading cause of disability worldwide, and accurate severity assessment using the Kellgren Lawrence (KL) grading system is critical for clinical decision making. However, radiographic distinctions between early disease stages, particularly KL1 and KL2, are subtle and frequently lead to inter-observer variability among radiologists. To address these challenges, we propose VLOrdinalFormer, a vision language guided ordinal learning framework for fully automated KOA grading from knee radiographs. The proposed method combines a ViT L16 backbone with CORAL based ordinal regression and a Contrastive Language Image Pretraining (CLIP) driven semantic alignment module, allowing the model to incorporate clinically meaningful textual concepts related to joint space narrowing, osteophyte formation, and subchondral sclerosis. To improve robustness and mitigate overfitting, we employ stratified five fold cross validation, class aware re weighting to emphasize challenging intermediate grades, and test time augmentation with global threshold optimization. Experiments conducted on the publicly available OAI kneeKL224 dataset demonstrate that VLOrdinalFormer achieves state of the art performance, outperforming CNN and ViT baselines in terms of macro F1 score and overall accuracy. Notably, the proposed framework yields substantial performance gains for KL1 and KL2 without compromising classification accuracy for mild or severe cases. In addition, interpretability analyses using Grad CAM and CLIP similarity maps confirm that the model consistently attends to clinically relevant anatomical regions. These results highlight the potential of vision language aligned ordinal transformers as reliable and interpretable tools for KOA grading and disease progression assessment in routine radiological practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。