arXiv:2605.12031cs.LGcs.CV2026-05

提出可应对模态缺失的多模态医学图像与表格数据联合学习框架。

Resilient Vision-Tabular Multimodal Learning under Modality Missingness

论文配图:Resilient Vision-Tabular Multimodal Learning under Modality Missingness
图 1 · 摘自论文原文
  • 通过可学习模态令牌与掩码自注意力实现中间融合,排除缺失模态影响。
  • 在MIMIC-CXR和MIMIC-IV数据集上,14类诊断任务表现优于基线,鲁棒性更强。
  • 适合临床真实场景中模态缺失严重的医疗多模态分析应用。

多模态深度学习在整合医学影像与结构化临床变量等异构数据方面展现出巨大潜力。然而,现有方法通常隐含假设模态数据完整,这在真实临床环境中极少成立,因整体模态或单个特征常缺失。本文提出一种专为普遍模态缺失设计的多模态变压器框架,用于视觉与表格数据的联合学习,无需依赖插补或启发式模型切换。该架构包含视觉、表格与多模态融合编码器三部分。单模态表示通过可学习模态令牌加权,并通过掩码自注意力实现中间融合,排除缺失令牌与模态的信息聚合与梯度传播。为进一步增强韧性,引入模态丢弃正则化策略,在训练中随机移除可用模态,促使模型在部分数据下利用互补信息。我们在配对了MIMIC-IV结构化数据的MIMIC-CXR数据集上,针对14种诊断发现的多标签分类任务进行评估。通过两种并行系统性压力测试协议,分别逐步增加每类模态的训练与推理缺失程度,覆盖从全多模态到全单模态的各种场景。在所有缺失程度下,所提方法均持续优于代表性基线,表现出更平滑的性能下降与更强鲁棒性。消融实验进一步证明,注意力级掩码与联合微调的中间融合是实现韧性多模态推理的关键。

原文摘要 · Abstract (English)

Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and structured clinical variables. However, most existing approaches implicitly assume complete modality availability, an assumption that rarely holds in real-world clinical settings where entire modalities and individual features are frequently missing. In this work, we propose a multimodal transformer framework for joint vision-tabular learning explicitly designed to operate under pervasive modality missingness, without relying on imputation or heuristic model switching. The architecture integrates three components: a vision, a tabular, and a multimodal fusion encoder. Unimodal representations are weighted through learnable modality tokens and fused via intermediate fusion with masked self-attention, which excludes missing tokens and modalities from information aggregation and gradient propagation. To further enhance resilience, we introduce a modality-dropout regularization strategy that stochastically removes available modalities during training, encouraging the model to exploit complementary information under partial data availability. We evaluate our approach on the MIMIC-CXR dataset paired with structured clinical data from MIMIC-IV for multilabel classification of 14 diagnostic findings with incomplete annotations. Two parallel systematic stress-test protocols progressively increase training and inference missingness in each modality separately, spanning fully multimodal to fully unimodal scenarios. Across all missingness regimes, the proposed method consistently outperforms representative baselines, showing smoother performance degradation and improved robustness. Ablation studies further demonstrate that attention-level masking and intermediate fusion with joint fine-tuning are key to resilient multimodal inference.

多模态学习医学影像模态缺失视觉表格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。