arXiv:2510.02922cs.CVcs.AI2025-10

用大模型融合超声与临床数据,提升颈动脉斑块风险评估精度

Multimodal Carotid Risk Stratification with Large Vision-Language Models: Benchmarking, Fine-Tuning, and Clinical Insights

  • 构建问答式框架,测试多种大视觉语言模型在多模态分析中的表现
  • 经低秩适配后,LLaVa-NeXT-Vicuna 在卒中风险分层上显著优于原始模型
  • 结合文本化临床数据可提升准确率,适合临床辅助决策场景

可靠的颈动脉粥样硬化疾病风险评估仍是临床重大挑战,需整合多样化的临床与影像信息,且结果应透明可解释。本研究探讨了先进大视觉语言模型(LVLMs)在融合超声影像(USI)与结构化临床、人口统计、实验室及蛋白生物标志物数据方面的潜力。提出一种模拟真实诊断流程的问答式评估框架,对比多种开源LVLM,包括通用型与医疗优化型模型。零样本实验显示,尽管模型能力强大,但并非所有模型能准确识别成像模态与解剖结构,且均在风险分类上表现不佳。通过低秩适配(LoRA)将LLaVa-NeXT-Vicuna适配至超声领域,显著提升卒中风险分层性能。进一步以文本形式融入多模态表格数据,增强了特异性和平衡准确率,达到与先前在相同数据集上训练的卷积神经网络(CNN)基线相当的水平。研究揭示了LVLM在基于超声的心血管风险预测中的潜力与局限,强调多模态融合、模型校准与领域适配对临床转化的重要性。

原文摘要 · Abstract (English)

Reliable risk assessment for carotid atheromatous disease remains a major clinical challenge, as it requires integrating diverse clinical and imaging information in a manner that is transparent and interpretable to clinicians. This study investigates the potential of state-of-the-art and recent large vision-language models (LVLMs) for multimodal carotid plaque assessment by integrating ultrasound imaging (USI) with structured clinical, demographic, laboratory, and protein biomarker data. A framework that simulates realistic diagnostic scenarios through interview-style question sequences is proposed, comparing a range of open-source LVLMs, including both general-purpose and medically tuned models. Zero-shot experiments reveal that even if they are very powerful, not all LVLMs can accurately identify imaging modality and anatomy, while all of them perform poorly in accurate risk classification. To address this limitation, LLaVa-NeXT-Vicuna is adapted to the ultrasound domain using low-rank adaptation (LoRA), resulting in substantial improvements in stroke risk stratification. The integration of multimodal tabular data in the form of text further enhances specificity and balanced accuracy, yielding competitive performance compared to prior convolutional neural network (CNN) baselines trained on the same dataset. Our findings highlight both the promise and limitations of LVLMs in ultrasound-based cardiovascular risk prediction, underscoring the importance of multimodal integration, model calibration, and domain adaptation for clinical translation.

多模态医学影像大模型风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。