arXiv:2509.03214cs.CV2025-09中稿 · BIBM 2025被引 4

用脑区文本生成+多模态融合提升脑疾病诊断准确率

RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion

  • 根据脑区激活和连接生成可复现的文本描述
  • 在ADHD-200和ABIDE数据集上准确率显著优于现有方法
  • 适合关注脑影像与文本联合分析的研究者

功能磁共振成像(fMRI)是探测脑功能的强大工具,但临床诊断受限于信噪比低、个体差异大,以及现有基于CNN和Transformer的模型对频率信息感知不足。此外,多数fMRI数据缺乏文本标注,难以解释区域激活与连接模式。本文提出RTGMFF框架,通过脑区驱动的文本生成与多模态特征融合实现脑疾病诊断。该框架包含三部分:(i) 脑区驱动的fMRI文本生成模块,将每个受试者的激活、连接、年龄、性别等信息凝练为可复现的文本标记;(ii) 混合频域-空间编码器,融合分层小波-Mamba分支与跨尺度Transformer编码器,捕捉频域结构与长程空间依赖;(iii) 自适应语义对齐模块,在共享空间中嵌入文本标记序列与视觉特征,使用正则化余弦相似性损失缩小模态差距。在ADHD-200和ABIDE基准上的实验表明,RTGMFF在诊断准确率、敏感性、特异性及ROC曲线下面积方面均超越现有方法。代码已开源。

原文摘要 · Abstract (English)

Functional magnetic resonance imaging (fMRI) is a powerful tool for probing brain function, yet reliable clinical diagnosis is hampered by low signal-to-noise ratios, inter-subject variability, and the limited frequency awareness of prevailing CNN- and Transformer-based models. Moreover, most fMRI datasets lack textual annotations that could contextualize regional activation and connectivity patterns. We introduce RTGMFF, a framework that unifies automatic ROI-level text generation with multimodal feature fusion for brain-disorder diagnosis. RTGMFF consists of three components: (i) ROI-driven fMRI text generation deterministically condenses each subject's activation, connectivity, age, and sex into reproducible text tokens; (ii) Hybrid frequency-spatial encoder fuses a hierarchical wavelet-mamba branch with a cross-scale Transformer encoder to capture frequency-domain structure alongside long-range spatial dependencies; and (iii) Adaptive semantic alignment module embeds the ROI token sequence and visual features in a shared space, using a regularized cosine-similarity loss to narrow the modality gap. Extensive experiments on the ADHD-200 and ABIDE benchmarks show that RTGMFF surpasses current methods in diagnostic accuracy, achieving notable gains in sensitivity, specificity, and area under the ROC curve. Code is available at https://github.com/BeistMedAI/RTGMFF.

脑影像分析多模态融合文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。