arXiv:2512.14026cs.CV2025-12AAAI被引 5

突破表格数据壁垒,让医学影像与表格数据联合自监督学习更通用高效

Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers

  • 用列名作语义线索,增强跨表格数据的表征学习能力
  • 在3个数据集上共4461名患者中,性能超越现有最佳方法
  • 适合需要融合多源异构医疗数据的研究者使用

近年来,融合医学影像与表格数据的多模态学习显著推动了临床决策。自监督学习(SSL)作为预训练大规模无标签图像-表格数据的强大范式,旨在学习具有区分性的表征。然而,现有图像-表格表示学习的自监督方法常受限于特定数据队列,主要源于其在处理异构表格数据时的僵化建模机制。这种跨表格障碍阻碍了多模态自监督方法在不同队列间有效学习可迁移的医学知识。本文提出一种新型自监督框架CITab,旨在以跨表格方式学习强大的多模态特征表示。我们从语义感知角度设计表格建模机制,将列标题作为语义线索,促进可迁移知识学习,并提升多数据源预训练的可扩展性。此外,我们提出原型引导的线性层混合(P-MoLin)模块,实现表格特征专业化,使模型能有效应对表格数据的异质性并挖掘潜在医学概念。我们在包含4461名受试者的三个公开数据队列上对阿尔茨海默病诊断任务进行了全面评估,实验结果表明,CITab优于当前最先进方法,为高效且可扩展的跨表格多模态学习开辟了新路径。

原文摘要 · Abstract (English)

Multi-modal learning integrating medical images and tabular data has significantly advanced clinical decision-making in recent years. Self-Supervised Learning (SSL) has emerged as a powerful paradigm for pretraining these models on large-scale unlabeled image-tabular data, aiming to learn discriminative representations. However, existing SSL methods for image-tabular representation learning are often confined to specific data cohorts, mainly due to their rigid tabular modeling mechanisms when modeling heterogeneous tabular data. This inter-tabular barrier hinders the multi-modal SSL methods from effectively learning transferrable medical knowledge shared across diverse cohorts. In this paper, we propose a novel SSL framework, namely CITab, designed to learn powerful multi-modal feature representations in a cross-tabular manner. We design the tabular modeling mechanism from a semantic-awareness perspective by integrating column headers as semantic cues, which facilitates transferrable knowledge learning and the scalability in utilizing multiple data sources for pretraining. Additionally, we propose a prototype-guided mixture-of-linear layer (P-MoLin) module for tabular feature specialization, empowering the model to effectively handle the heterogeneity of tabular data and explore the underlying medical concepts. We conduct comprehensive evaluations on Alzheimer's disease diagnosis task across three publicly available data cohorts containing 4,461 subjects. Experimental results demonstrate that CITab outperforms state-of-the-art approaches, paving the way for effective and scalable cross-tabular multi-modal learning.

多模态学习自监督学习医学影像表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。