用可解释的视觉语言模型提升腰椎管狭窄诊断准确率
An Explainable Vision-Language Model Framework with Adaptive PID-Tversky Loss for Lumbar Spinal Stenosis Diagnosis
- 通过空间块交叉注意力实现病变精准定位
- 自适应PID-Tversky损失使少数类分割准确率达95.12%
- 生成放射科报告,适合临床医生辅助诊断
腰椎管狭窄(LSS)诊断依赖人工解读多视角磁共振成像,存在主观差异大、耗时长的问题。现有视觉语言模型难以应对临床分割数据集中的极端类别不平衡,且因全局池化破坏解剖层级而影响空间精度。本文提出端到端可解释视觉语言模型框架,通过空间块交叉注意力模块实现文本引导的精确病灶定位;设计基于控制理论的自适应PID-Tversky损失函数,动态调整训练惩罚以优化难分少数类。结合基础视觉语言模型与自动报告生成模块,该框架在诊断分类上达到90.69%准确率,分割宏平均Dice分数达0.9512,报告生成CIDEr评分为92.80%。同时可将复杂分割结果转化为放射科风格报告,实现透明可解释的AI辅助诊断,兼顾人类监督与诊断效能。
原文摘要 · Abstract (English)
Lumbar Spinal Stenosis (LSS) diagnosis remains a critical clinical challenge, with diagnosis heavily dependent on labor-intensive manual interpretation of multi-view Magnetic Resonance Imaging (MRI), leading to substantial inter-observer variability and diagnostic delays. Existing vision-language models simultaneously fail to address the extreme class imbalance prevalent in clinical segmentation datasets while preserving spatial accuracy, primarily due to global pooling mechanisms that discard crucial anatomical hierarchies. We present an end-to-end Explainable Vision-Language Model framework designed to overcome these limitations, achieved through two principal objectives. We propose a Spatial Patch Cross-Attention module that enables precise, text-directed localization of spinal anomalies with spatial precision. A novel Adaptive PID-Tversky Loss function by integrating control theory principles dynamically further modifies training penalties to specifically address difficult, under-segmented minority instances. By incorporating foundational VLMs alongside an Automated Radiology Report Generation module, our framework demonstrates considerable performance: a diagnostic classification accuracy of 90.69%, a macro-averaged Dice score of 0.9512 for segmentation, and a CIDEr score of 92.80%. Furthermore, the framework shows explainability by converting complex segmentation predictions into radiologist-style clinical reports, thereby establishing a new benchmark for transparent, interpretable AI in clinical medical imaging that keeps essential human supervision while enhancing diagnostic capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。