用视觉语言模型自动评估心脏磁共振图像质量,提升消融手术规划安全性。
Toward Vision Language Model-based Assessment of Clinical Quality and Usability of LGE-MR Images for Cardiac Ablation Planning

- 分两阶段:先生成放射科风格报告,再推理出是否可用
- 在20例患者60张图像上,临床可用性判断准确率达100%
- 首次实现可解释的临床级图像质量评估,适合医学影像团队使用
LGE心脏磁共振广泛用于房颤患者左心房纤维化评估和消融规划,因纤维组织区域识别对导管消融至关重要。但图像质量差常导致消融靶点误定位,直接影响手术安全与效果。目前图像是否满足最低质量标准仍由放射科医生主观判断,无自动化系统记录,却是图像质量评估中最关键的安全输出。噪声、运动伪影及边界不清等质量问题显著影响后续分割与临床决策可靠性。专家手动评估主观性强且难扩展,现有自动化方法仅输出数值评分,缺乏可解释的临床逻辑。本文提出一种两阶段视觉语言模型(VLM)框架,用于左心房LGE-MRI的临床导向质量评估。第一阶段,微调VLM生成结构化放射科风格报告,预测五项放射科定义的质量标准:噪声、运动伪影、左心房边界准确性、肺静脉区准确性、下分割严重程度。第二阶段,基于GPT的推理模块将预测结果映射为结构化评分与二元临床可用性决策。我们构建了包含20名患者60对标注图像-文本的数据集,并对比四种先进VLM架构。InternVL2在准则层面表现最优(平均准确率ACC=0.65,皮尔逊相关系数PLCC=0.79),DeepSeek在临床可用性判断上达成完美一致(准确率Acc=1.00,kappa=1.00)。
原文摘要 · Abstract (English)
LGE cardiac MRI is widely used for left atrial fibrosis assessment and ablation planning in atrial fibrillation patients as knowledge of fibrotic tissue regions identified from LGE-MRI is critical for catheter ablation. Often, poor quality images used during ablation planning can cause mis-localization of ablation targets, directly impacting procedure safety and outcome. The decision of whether a scan meets the minimum quality threshold for ablation planning is currently made informally by the reviewing radiologist and is not captured by any automated system, yet it is arguably the most safety-critical output of the image quality assessment (IQA) process. However, variations in image quality caused by noise, motion artifacts, and poor boundary definition significantly compromise the reliability of downstream segmentation and clinical decision-making tasks. Manual quality assessment by expert radiologists is subjective and difficult to scale, while existing automated methods produce scalar scores without interpretable clinical reasoning. In this work, we propose a two-stage vision language model (VLM) framework for clinically grounded image quality assessment of left atrial LGE-MRI. In the first stage, a fine-tuned VLM generates structured radiology-style quality reports predicting five radiologist-defined criteria: Noise, Motion Artifact, LA Boundary Accuracy, PV Region Accuracy, and Under-segmentation Severity. In the second stage, a GPT-based reasoning module maps the predicted quality and reports to a structured quality scores and binary clinical usability decision for ablation planning. We curate a dataset of 60 annotated image slice-text pairs from 20 patients and benchmark four state-of-the-art VLM architectures. InternVL2 achieves the highest criterion-level accuracy (Avg ACC=0.65, PLCC=0.79), while DeepSeek achieves perfect clinical usability agreement (Acc=1.00, kappa=1.00).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。