arXiv:2512.08462cs.LG2025-12

用Transformer融合脑影像与医疗元数据,提升脑状态解码准确率

Transformers for Multimodal Brain State Decoding: Integrating Functional Magnetic Resonance Imaging Data and Medical Metadata

  • 采用Transformer架构融合fMRI数据与DICOM元数据
  • 在多模态输入下显著提升解码精度与模型可解释性
  • 适合神经科学、临床诊断与个性化医疗研究者

从功能磁共振成像(fMRI)数据中解码脑状态对推动神经科学和临床应用至关重要。尽管传统机器学习与深度学习方法已在处理高维、复杂的fMRI数据方面取得进展,但往往未能充分利用数字医学成像与通信标准(DICOM)元数据提供的上下文信息。本文提出一种基于Transformer的新型多模态框架,整合fMRI数据与DICOM元数据。通过注意力机制,该方法能够捕捉复杂的时空模式与上下文关联,提升模型的准确性、可解释性与鲁棒性。该框架在临床诊断、认知神经科学及个性化医疗中具有广泛应用潜力。文中也讨论了元数据异质性与计算开销等局限性,并提出了优化可扩展性与泛化能力的未来方向。

原文摘要 · Abstract (English)

Decoding brain states from functional magnetic resonance imaging (fMRI) data is vital for advancing neuroscience and clinical applications. While traditional machine learning and deep learning approaches have made strides in leveraging the high-dimensional and complex nature of fMRI data, they often fail to utilize the contextual richness provided by Digital Imaging and Communications in Medicine (DICOM) metadata. This paper presents a novel framework integrating transformer-based architectures with multimodal inputs, including fMRI data and DICOM metadata. By employing attention mechanisms, the proposed method captures intricate spatial-temporal patterns and contextual relationships, enhancing model accuracy, interpretability, and robustness. The potential of this framework spans applications in clinical diagnostics, cognitive neuroscience, and personalized medicine. Limitations, such as metadata variability and computational demands, are addressed, and future directions for optimizing scalability and generalizability are discussed.

脑状态解码多模态学习Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。