arXiv:2505.04964cs.CV2025-05

用专业标注数据微调视觉语言模型,提升冠脉造影图像的临床报告生成能力。

CAG-VLM: Fine-Tuning of a Large-Scale Model to Recognize Angiographic Images for Next-Generation Diagnostic Systems

  • 构建双阶段标注流程,用1.4万张图像训练分类模型,准确率高达0.96
  • 基于专家验证数据微调三个开源模型,最佳表现达7.2分(满分10)
  • 专为冠脉造影设计,适合辅助心脏病医生生成诊断与治疗建议

冠状动脉造影(CAG)是评估冠心病的金标准,但其解读和治疗方案制定高度依赖心脏科专家。为实现AI辅助决策,我们提出一个两阶段、由医师标注的流程,并构建了一个日英双语的CAG图像-报告数据集。首先,从539例检查中抽取14,686帧图像,进行关键帧检测与左右侧别标注;基于此数据训练的ConvNeXt-Base CNN在侧别分类上达到0.96的F1值,即使在低对比度图像上也表现良好。其次,将该CNN应用于243例独立检查,提取1,114个关键帧,并配对术前报告及专家验证的诊断与治疗摘要,形成平行语料库。随后,使用LoRA对三个开源视觉语言模型(PaliGemma2、Gemma3、ConceptCLIP-enhanced Gemma3)进行微调,并通过VLScore与心脏科医生评审评估。尽管带LoRA的PaliGemma2在VLScore上最高,但带LoRA的Gemma3获得最高医生评分(平均7.20/10),被命名为CAG-VLM。结果表明,经专业化微调的视觉语言模型可有效辅助心脏科医生从CAG图像生成临床报告与治疗建议。

原文摘要 · Abstract (English)

Coronary angiography (CAG) is the gold-standard imaging modality for evaluating coronary artery disease, but its interpretation and subsequent treatment planning rely heavily on expert cardiologists. To enable AI-based decision support, we introduce a two-stage, physician-curated pipeline and a bilingual (Japanese/English) CAG image-report dataset. First, we sample 14,686 frames from 539 exams and annotate them for key-frame detection and left/right laterality; a ConvNeXt-Base CNN trained on this data achieves 0.96 F1 on laterality classification, even on low-contrast frames. Second, we apply the CNN to 243 independent exams, extract 1,114 key frames, and pair each with its pre-procedure report and expert-validated diagnostic and treatment summary, yielding a parallel corpus. We then fine-tune three open-source VLMs (PaliGemma2, Gemma3, and ConceptCLIP-enhanced Gemma3) via LoRA and evaluate them using VLScore and cardiologist review. Although PaliGemma2 w/LoRA attains the highest VLScore, Gemma3 w/LoRA achieves the top clinician rating (mean 7.20/10); we designate this best-performing model as CAG-VLM. These results demonstrate that specialized, fine-tuned VLMs can effectively assist cardiologists in generating clinical reports and treatment recommendations from CAG images.

医学影像视觉语言模型冠脉造影临床辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。