arXiv:2512.11060cs.CV2025-12

用合成血管和病灶数据训练视觉语言模型,提升医学影像推理能力

Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning

  • 通过可控生成真实眼底血管与糖尿病视网膜病变特征,自动配文本
  • 在10万张合成OCTA图像上训练,零样本分类准确率达89.67%
  • 显著提升临床解释质量与病灶定位,适合医疗AI可解释性研究

视觉-语言模型(VLM)为可解释医学诊断提供了可能,使用户可在不同模态下询问临床解释。然而,实现精细推理需大规模图文数据集。在光学相干断层扫描血管成像(OCTA)等专业领域,带有病灶精准描述的文本数据稀缺甚至缺失。为此,本文提出合成血管推理框架(SVR),可控生成包含糖尿病视网膜病变特征(如毛细血管缺失、微动脉瘤、新生血管、迂曲)的真实视网膜血管图像,并自动生成细粒度推理文本。基于此构建了含10万对数据的OCTA-100K-SVR数据集。实验表明,在该数据集上训练的通用VLM(Qwen3-VL-8b)在真实OCTA图像上达到89.67%的零样本平衡分类准确率,优于监督基线。专家评估进一步证实其显著提升了临床数据中的解释质量与病灶定位能力。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) offer a promising path toward interpretable medical diagnosis by allowing users to ask about clinical explanations alongside predictions and across different modalities. However, training VLMs for detailed reasoning requires large-scale image-text datasets. In many specialized domains, for example in reading Optical Coherence Tomography Angiography (OCTA) images, such precise text with grounded description of pathologies is scarce or even non-existent. To overcome this bottleneck, we introduce Synthetic Vasculature Reasoning (SVR), a framework that controllably synthesizes images and corresponding text, specifically: realistic retinal vasculature with Diabetic Retinopathy (DR) features: capillary dropout, microaneurysms, neovascularization, and tortuosity, while automatically generating granular reasoning texts. Based on this we curate OCTA-100K-SVR, an OCTA image-reasoning dataset with 100,000 pairs. Our experiments show that a general-purpose VLM (Qwen3-VL-8b) trained on the dataset achieves a zero-shot balanced classification accuracy of 89.67% on real OCTA images, outperforming supervised baselines. Through human expert evaluation we also demonstrate that it significantly enhances explanation quality and pathology localization on clinical data.

视觉语言模型医学影像可解释性数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。