arXiv:2506.19439cs.CV2025-06被引 2

解决医学影像与表格数据融合难题,提升小样本下的诊断准确率。

AMF-MedIT: An Efficient Align-Modulation-Fusion Framework for Medical Image-Tabular Data

  • 设计自监督对齐调制融合模块,动态平衡模态贡献并处理维度差异。
  • 在低数据场景下分类准确率显著提升,噪声环境下仍保持稳定性能。
  • 适合临床人工智能应用,尤其适用于数据稀缺的医疗场景。

多模态医学分析结合影像与表格数据日益受到关注。然而,跨模态特征维度不一致、模态贡献差异及高维表格数据噪声等问题使有效融合仍具挑战。为此,我们提出AMF-MedIT框架,一种面向医学图像与表格数据集成的高效对齐-调制-融合方法,特别适用于数据稀缺条件。基于自监督学习策略,引入自适应调制与融合(AMF)模块,一种新型轻量级融合范式,可协调维度差异并动态平衡模态贡献。该模块融合先验知识以指导模态分配,并使用特征掩码结合幅度与泄漏损失来调节单模态特征的维度与幅值。此外,我们设计了FT-Mamba,一种利用选择性机制高效处理噪声医学表格数据的强大编码器。大量实验(包括临床噪声模拟)表明,AMF-MedIT在多模态分类任务中实现了更高的准确率、鲁棒性和数据效率。可解释性分析进一步揭示了FT-Mamba如何塑造多模态预训练并增强图像编码器注意力,凸显该框架在可靠高效临床人工智能应用中的实际价值。

原文摘要 · Abstract (English)

Multimodal medical analysis combining image and tabular data has gained increasing attention. However, effective fusion remains challenging due to cross-modal discrepancies in feature dimensions and modality contributions, as well as the noise from high-dimensional tabular inputs. To address these problems, we present AMF-MedIT, an efficient Align-Modulation-Fusion framework for medical image and tabular data integration, particularly under data-scarce conditions. Built upon a self-supervised learning strategy, we introduce the Adaptive Modulation and Fusion (AMF) module, a novel, streamlined fusion paradigm that harmonizes dimension discrepancies and dynamically balances modality contributions. It integrates prior knowledge to guide the allocation of modality contributions in the fusion and employs feature masks together with magnitude and leakage losses to adjust the dimensionality and magnitude of unimodal features. Additionally, we develop FT-Mamba, a powerful tabular encoder leveraging a selective mechanism to handle noisy medical tabular data efficiently. Extensive experiments, including simulations of clinical noise, demonstrate that AMF-MedIT achieves superior accuracy, robustness, and data efficiency across multimodal classification tasks. Interpretability analyses further reveal how FT-Mamba shapes multimodal pretraining and enhances the image encoder's attention, highlighting the practical value of our framework for reliable and efficient clinical artificial intelligence applications.

医学多模态数据融合小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。