arXiv:2605.25968cs.CV2026-05

解决医学影像与表格数据缺失问题,提升诊断鲁棒性。

Context-driven Missing-Modality Learning for Robust Medical Diagnosis with Image-Tabular Data

论文配图:Context-driven Missing-Modality Learning for Robust Medical Diagnosis with Image-Tabular Data
图 1 · 摘自论文原文
  • 用上下文驱动的渐进式合成与对齐方法重建缺失模态
  • 在三个数据集上平均提升AUC 1.26%~1.32%
  • 适合处理临床中模态不完整场景的医疗诊断任务

多模态数据融合影像与临床表格记录对精准医疗诊断至关重要,但临床实践中特定模态的任意缺失普遍存在,严重削弱多模态模型性能。现有方法或直接丢弃缺失模态导致信息损失,或难以捕捉复杂模态间依赖关系而合成失败。为此,本文提出上下文驱动的缺失模态学习(CMML)框架,通过序列化模态合成与语义对齐,在任意缺失条件下实现鲁棒诊断。具体地,设计基于级联残差变换器的自编码器(CRTA),利用可学习的上下文标记作为数据级语义先验,捕获模态间依赖并合成关键缺失表示;进一步通过模态专属记忆库增强表示。为缓解原始可用与合成表示间的差异,将学习到的上下文标记转化为实例自适应语义参考,融合CRTA输出的多模态表示。该参考引导异构模态表示对齐至统一空间,并最终应用类感知对比精炼以挖掘判别性诊断线索。在皮肤病变(Derm7pt)、眼部疾病(ODIR)和脑膜瘤(MEN)数据集上的大量实验表明,CMML显著优于现有最先进方法,平均AUC分别提升1.26%、0.97%和1.32%。

原文摘要 · Abstract (English)

While multimodal data integrating diverse imaging and clinical tabular records is crucial for accurate medical diagnosis, the arbitrary absence of specific modalities is prevalent in clinical practice, severely degrading the performance of multimodal models. Existing methods either discard missing modalities, leading to information loss, or struggle to synthesize them without capturing complex inter-modal dependencies. To address these limitations, we propose a novel Context-driven Missing-Modality Learning (CMML) framework, which sequentially performs modality synthesis and semantic alignment to achieve robust diagnosis under arbitrary missing conditions. Specifically, we design a Cascade Residual Transformer-based Autoencoder (CRTA) that leverages learnable context tokens acting as dataset-level semantic prior to capture inter-modal dependencies and synthesize key missing representations. These representations are further enriched by modality-specific memory banks. To resolve the discrepancy between original available and synthesized representations, we transform the learned context tokens into instance-adaptive semantic references by infusing multimodal representations from the CRTA's outputs. This reference guides the alignment of heterogeneous modality representations into a unified space, where class-aware contrastive refinement is finally applied to explore discriminative diagnostic cues. Extensive evaluations on skin lesion (Derm7pt), ocular disease (ODIR), and meningioma (MEN) datasets demonstrate that CMML significantly outperforms state-of-the-art (SOTA) methods, yielding AVG AUC improvements of 1.26%, 0.97%, and 1.32%, respectively.

医学诊断多模态学习缺失数据图像表格融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。