用视觉引导文本融合,提升脑卒中预后预测准确率
Vision-Core Guided Contrastive Learning for Balanced Multi-modal Prognosis Prediction of Stroke

- 用大模型从MRI生成结构化文本,增强多模态表示
- 视觉特征引导文本融合,实现跨模态深度交互
- 在真实临床数据上表现领先,适合医疗AI研究者
深度学习与多模态融合在医学诊断中展现出变革潜力,但缺血性脑卒中的精准预后仍面临挑战。现有方法多局限于双模态融合,缺乏对医学影像、结构化临床数据与非结构化文本的三模态整合框架;且常无法建立模态间的深层双向交互。为此,本文提出一种新型三模态融合模型。首先,利用大语言模型(LLM)从脑部MRI自动生成半结构化诊断文本,缓解专家标注稀缺问题,并作为语义正则化增强多模态融合鲁棒性。其次,设计视觉条件双重对齐融合模块(VDAFM),以视觉特征为条件先验,指导细粒度文本交互,通过双重语义对齐损失实现动态深度融合,有效缓解模态异质性。在真实临床数据集上的大量实验表明,该模型达到当前最优性能。
原文摘要 · Abstract (English)
Deep learning and multi-modal fusion have demonstrated transformative potential in medical diagnosis by integrating diverse data sources. However, accurate prognosis for ischemic stroke remains challenging due to limitations in existing multi-modal approaches. First, current methods are predominantly confined to dual-modal fusion, lacking a framework that effectively integrates the trifecta of medical images, structured clinical data, and unstructured text. Second, they often fail to establish deep bidirectional interactions between modalities; To address these critical gaps, this paper proposes a novel tri-modal fusion model for ischemic stroke prognosis. Our approach first enriches the data representation by employing a Large Language Model (LLM) to automatically generate semi-structured diagnostic text from brain MRIs. This process not only addresses the scarcity of expert annotations but also serves as a regularized semantic enhancement, improving multimodal fusion robustness. Furthermore, we design a core component termed the Vision-Conditioned Dual Alignment Fusion Module (VDAFM), which strategically uses visual features as a conditional prior to guide fine-grained interaction with the generated text. This module achieves a dynamic and profound fusion through a dual semantic alignment loss, effectively mitigating modal heterogeneity. Extensive experiments on a real-world clinical dataset demonstrate that our model achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。