用对比学习对齐腰椎MRI和报告,提升疼痛诊断准确率
Revolutionizing Precise Low Back Pain Diagnosis via Contrastive Learning
- 通过对比学习将MRI图像与放射报告映射到同一向量空间
- 在不平衡数据下实现95%准确率和94.75%F1分数
- 适合临床辅助诊断与骨骼肌肉系统自动化分析
腰椎疼痛影响全球数百万人群,亟需能够联合分析复杂医学影像与文本报告的鲁棒诊断模型。本文提出LumbarCLIP,一种基于对比语言-图像预训练的多模态框架,将腰椎MRI扫描与对应的放射科描述对齐。该模型基于包含轴向MRI视图与专家撰写报告的自建数据集,融合视觉编码器(ResNet-50、Vision Transformer、Swin Transformer)与BERT文本编码器,提取密集表征,并通过可学习投影头(线性或非线性)映射至共享嵌入空间,采用软CLIP损失进行归一化对比训练。在下游分类任务中,模型在测试集上达到最高95.00%准确率与94.75% F1分数,即使存在固有类别不平衡。大量消融实验表明,线性投影头比非线性变体更有效实现跨模态对齐。LumbarCLIP为自动化肌肉骨骼疾病诊断与临床决策支持提供了有力基础。
原文摘要 · Abstract (English)
Low back pain affects millions worldwide, driving the need for robust diagnostic models that can jointly analyze complex medical images and accompanying text reports. We present LumbarCLIP, a novel multimodal framework that leverages contrastive language-image pretraining to align lumbar spine MRI scans with corresponding radiological descriptions. Built upon a curated dataset containing axial MRI views paired with expert-written reports, LumbarCLIP integrates vision encoders (ResNet-50, Vision Transformer, Swin Transformer) with a BERT-based text encoder to extract dense representations. These are projected into a shared embedding space via learnable projection heads, configurable as linear or non-linear, and normalized to facilitate stable contrastive training using a soft CLIP loss. Our model achieves state-of-the-art performance on downstream classification, reaching up to 95.00% accuracy and 94.75% F1-score on the test set, despite inherent class imbalance. Extensive ablation studies demonstrate that linear projection heads yield more effective cross-modal alignment than non-linear variants. LumbarCLIP offers a promising foundation for automated musculoskeletal diagnosis and clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。