arXiv:2510.01899cs.LGcs.AI2025-10被引 3

融合多源医疗数据的模型,提升早期疾病检测能力

Multimodal Foundation Models for Early Disease Detection

  • 用跨模态注意力融合电子病历、影像等数据
  • 在多种疾病场景下实现稳定准确的早期诊断
  • 支持缺失数据,适合临床实际应用

医疗数据涵盖电子病历(EHR)、医学影像、基因组和可穿戴设备信号,但现有诊断模型通常孤立处理各模态,难以捕捉早期跨模态疾病特征。本文提出一种基于Transformer的多模态基础模型,通过模态特定编码器与跨模态注意力机制,将异构临床数据映射至共享潜在空间,并利用多头注意力与残差归一化进行融合。模型在模拟早期疾病模式的数据集上训练,涵盖EHR序列、影像块、基因组谱和可穿戴信号,包含缺失模态和标签噪声情形。采用监督分类结合自监督重建与对比对齐策略,增强鲁棒性。实验表明,该模型在早期检测任务中表现优异,具备稳定的分类指标、可靠的不确定性估计及可解释的注意力模式。方法支持灵活预训练-微调,适用于肿瘤学、心脏病学和神经病学中的精准诊疗,能处理不完整输入,推动早期疾病检测发展。

原文摘要 · Abstract (English)

Healthcare data now span EHRs, medical imaging, genomics, and wearable sensors, but most diagnostic models still process these modalities in isolation. This limits their ability to capture early, cross-modal disease signatures. This paper introduces a multimodal foundation model built on a transformer architecture that integrates heterogeneous clinical data through modality-specific encoders and cross-modal attention. Each modality is mapped into a shared latent space and fused using multi-head attention with residual normalization. We implement the framework using a multimodal dataset that simulates early-stage disease patterns across EHR sequences, imaging patches, genomic profiles, and wearable signals, including missing-modality scenarios and label noise. The model is trained using supervised classification together with self-supervised reconstruction and contrastive alignment to improve robustness. Experimental evaluation demonstrates strong performance in early-detection settings, with stable classification metrics, reliable uncertainty estimates, and interpretable attention patterns. The approach moves toward a flexible, pretrain-and-fine-tune foundation model that supports precision diagnostics, handles incomplete inputs, and improves early disease detection across oncology, cardiology, and neurology applications.

多模态早期诊断医疗AI基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。