arXiv:2510.20736cs.LG2025-10NeurIPS

用狄利克雷过程自动强化多模态中关键特征,提升融合效果

Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet Process

  • 基于狄利克雷过程动态分配特征权重,突出重要特征
  • 在多个数据集上优于现有方法,跨模态对齐更优
  • 适合需要精准特征选择的医疗、金融等多模态场景

在医疗、金融等实际场景中,有效的多模态融合方法日益重要。核心挑战在于如何在保持各模态特征表达力的同时学习跨模态交互。以往方法主要关注模态间的对齐,但过度强调边际分布对齐会引入过多正则化,抑制各模态中的有意义表示。狄利克雷过程(DP)混合模型是一种强大的贝叶斯非参数方法,其‘富者愈富’特性可自动放大最显著的特征。受此启发,我们提出一种新型DP驱动的多模态学习框架,能自动平衡显著的模态内表示学习与跨模态对齐。具体地,假设每种模态服从多元高斯混合分布,并采用DP计算各成分的混合权重。该范式使DP能够动态分配特征贡献并选择最显著的特征,利用其‘富者愈富’特性,促进多模态特征融合。在多个多模态数据集上的大量实验表明,该模型性能优于其他对比方法。消融分析进一步验证了DP在对齐模态分布方面的有效性及其对关键超参数变化的鲁棒性。代码已匿名发布于 https://github.com/HKU-MedAI/DPMM.git。

原文摘要 · Abstract (English)

Developing effective multimodal fusion approaches has become increasingly essential in many real-world scenarios, such as health care and finance. The key challenge is how to preserve the feature expressiveness in each modality while learning cross-modal interactions. Previous approaches primarily focus on the cross-modal alignment, while over-emphasis on the alignment of marginal distributions of modalities may impose excess regularization and obstruct meaningful representations within each modality. The Dirichlet process (DP) mixture model is a powerful Bayesian non-parametric method that can amplify the most prominent features by its richer-gets-richer property, which allocates increasing weights to them. Inspired by this unique characteristic of DP, we propose a new DP-driven multimodal learning framework that automatically achieves an optimal balance between prominent intra-modal representation learning and cross-modal alignment. Specifically, we assume that each modality follows a mixture of multivariate Gaussian distributions and further adopt DP to calculate the mixture weights for all the components. This paradigm allows DP to dynamically allocate the contributions of features and select the most prominent ones, leveraging its richer-gets-richer property, thus facilitating multimodal feature fusion. Extensive experiments on several multimodal datasets demonstrate the superior performance of our model over other competitors. Ablation analysis further validates the effectiveness of DP in aligning modality distributions and its robustness to changes in key hyperparameters. Code is anonymously available at https://github.com/HKU-MedAI/DPMM.git

多模态学习狄利克雷过程特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。