用表格基础模型增强图像+表格多模态学习,抗缺失数据能力强
TIME: TabPFN-Integrated Multimodal Engine for Robust Tabular-Image Learning
- 用冻结的TabPFN生成抗缺失数据的表格嵌入,与图像特征融合
- 在完整和缺失表格数据上均超越基线,医学与自然数据集表现稳定
- 适合医疗等真实场景中存在大量缺失值的多模态任务
表格-图像多模态学习融合结构化表格数据与图像数据,在医疗等领域有广阔前景。但面临两大挑战:(1) 缺乏如视觉与语言领域那样的标准化预训练表格表示;(2) 表格数据中常见缺失值难以处理。为此,我们提出TabPFN集成多模态引擎(TIME),基于近期提出的表格基础模型TabPFN。TIME将TabPFN作为冻结的表格编码器,生成对缺失数据天然鲁棒的强嵌入,并与预训练视觉骨干网络提取的图像特征融合。我们探索多种融合策略与表格编码器,在自然与医学数据集上进行评估。大量实验表明,无论输入表格是否完整,TIME均持续优于对比基线,凸显其在真实多模态学习场景中的实用价值。
原文摘要 · Abstract (English)
Tabular-image multimodal learning, which integrates structured tabular data with imaging data, holds great promise for a variety of tasks, especially in medical applications. Yet, two key challenges remain: (1) the lack of a standardized, pretrained representation for tabular data, as is commonly available in vision and language domains; and (2) the difficulty of handling missing values in the tabular modality, which are common in real-world medical datasets. To address these issues, we propose the TabPFN-Integrated Multimodal Engine (TIME), a novel multimodal framework that builds on the recently introduced tabular foundation model, TabPFN. TIME leverages TabPFN as a frozen tabular encoder to generate robust, strong embeddings that are naturally resilient to missing data, and combines them with image features from pretrained vision backbones. We explore a range of fusion strategies and tabular encoders, and evaluate our approach on both natural and medical datasets. Extensive experiments demonstrate that TIME consistently outperforms competitive baselines across both complete and incomplete tabular inputs, underscoring its practical value in real-world multimodal learning scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。