arXiv:2505.17619cs.CV2025-05

用视觉语言模型评估生成血管造影质量,提升临床可用性

CAS-IQA: Teaching Vision-Language Models for Synthetic Angiography Quality Assessment

  • 基于视觉语言模型融合辅助图像信息,实现细粒度质量评分
  • 在3565张合成血管造影上,性能显著优于现有方法
  • 专为临床需求设计,适合医学影像质量评估研究者

现代生成模型生成的合成X射线血管造影有望减少介入手术中对比剂的使用,但低质量图像会显著增加操作风险,亟需可靠的图像质量评估(IQA)方法。现有IQA模型无法利用相关辅助图像作为参考,且缺乏针对临床任务的细粒度指标。为此,本文提出CAS-IQA,一种基于视觉语言模型(VLM)的框架,通过有效整合相关图像的辅助信息,预测细粒度质量分数。构建了无真实血管造影数据集的CAS-3K数据集,包含3,565张合成血管造影及其评分标注。为确保临床相关性,定义了三项任务特定评估指标。此外,设计了多路径特征融合与路由(MUST)模块,通过自适应融合和路由视觉令牌至特定指标分支,增强图像表征。在CAS-3K数据集上的大量实验表明,CAS-IQA显著优于当前最优IQA方法。

原文摘要 · Abstract (English)

Synthetic X-ray angiographies generated by modern generative models hold great potential to reduce the use of contrast agents in vascular interventional procedures. However, low-quality synthetic angiographies can significantly increase procedural risk, underscoring the need for reliable image quality assessment (IQA) methods. Existing IQA models, however, fail to leverage auxiliary images as references during evaluation and lack fine-grained, task-specific metrics necessary for clinical relevance. To address these limitations, this paper proposes CAS-IQA, a vision-language model (VLM)-based framework that predicts fine-grained quality scores by effectively incorporating auxiliary information from related images. In the absence of angiography datasets, CAS-3K is constructed, comprising 3,565 synthetic angiographies along with score annotations. To ensure clinically meaningful assessment, three task-specific evaluation metrics are defined. Furthermore, a Multi-path featUre fuSion and rouTing (MUST) module is designed to enhance image representations by adaptively fusing and routing visual tokens to metric-specific branches. Extensive experiments on the CAS-3K dataset demonstrate that CAS-IQA significantly outperforms state-of-the-art IQA methods by a considerable margin.

医学影像图像质量评估视觉语言模型生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。