基于数据自适应融合的多模态学习框架,提升预测准确性。
DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing

- 通过随机矩阵理论与非参数依赖度量,从数据中自动学习模态间融合结构。
- 在模拟与真实数据上均优于主流方法,实现更高精度的表达与生存预测。
- 适合处理异构多模态数据,尤其适用于信号弱、依赖复杂场景。
多模态监督学习旨在利用多种异构数据源提升预测性能。核心挑战在于确定模态间的融合粒度:过度融合会放大噪声,而融合不足则无法利用跨模态依赖关系。现有方法依赖预设的融合架构(如早期或晚期融合),难以适应模态间的实际依赖结构。本文提出DAIF,一种数据驱动的中间融合框架,结合随机矩阵理论与非参数依赖度量,直接从数据中学习融合结构。在贝叶斯多模态因子模型下,潜变量先验决定跨模态依赖。该方法基于估计的模态间依赖进行聚类,并对每簇执行经验贝叶斯先验估计。这些估计先验用于构建近似消息传递(AMP)框架中的去噪器,生成低维去噪特征,在相关模态间共享信息的同时保留模态特异性信号。最终嵌入用于下游监督预测。通过不同依赖结构与信号水平下的模拟实验,以及在两个真实多模态数据集上的验证——包括三模态TEA-seq数据集(Swanson et al., 2021)和TCGA-BRCA数据集(Goldman et al., 2020)——分别用于预测T细胞分化标志物表达水平与患者生存率,结果表明该方法在两项任务中均达到或超过当前最优水平,展现了其在多样监督学习任务中的通用性。
原文摘要 · Abstract (English)
Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance. A central challenge is determining the fusion granularity across modalities: over-integration may amplify noise while under-integration fails to exploit cross-modal dependence. Existing approaches rely on pre-specified fusion architectures, from early to late fusion, that may not adapt to the underlying dependence structure among modalities. We propose DAIF, a data adaptive intermediate fusion framework that combines random matrix theory and non-parametric dependence measures to learn fusion structure directly from data. We operate under a Bayesian multimodal factor model where the prior on the latent factors determines the cross-modal dependence. Our method clusters modalities based on estimated intermodal dependence, then performs clusterwise empirical Bayes estimation of the priors. These estimated priors are used to construct denoisers within an approximate message passing (AMP) framework, yielding denoised low-dimensional features that borrow strength across related modalities while preserving modality-specific signal. The resulting embeddings are used for downstream supervised prediction. We evaluate the framework through simulations under varying dependence structures and signal regimes, comparing against several benchmark methods, and demonstrate its practical utility on two multimodal datasets, namely a trimodal TEA-seq dataset (Swanson et al., 2021) and TCGA-BRCA dataset (Goldman et al., 2020). In the first example, we predict the expression level of a T-cell differentiation marker protein and in the second case we analyze patient survival prediction based on multimodal information. Our method competes with or outperforms the state-of-the-art techniques in both prediction problems, demonstrating its versatility across diverse supervised learning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。