arXiv:2606.02659cs.LGcs.AI2026-06AAAI

对比学习增强多模态融合,动态处理缺失数据。

CL-DMDF:Dynamic Multimodal Data Fusion Model Based on Contrastive Learning

论文配图:CL-DMDF:Dynamic Multimodal Data Fusion Model Based on Contrastive Learning
图 1 · 摘自论文原文
  • 跨特征与模态维度的注意力机制,精准捕捉重要信息。
  • 通过实体-中心对比学习提升特征区分能力。
  • 自适应融合模块支持动态任务,适合真实场景应用。

多模态数据融合旨在整合多源信息以揭示潜在关联与互补模式,从而提升数据处理与决策能力。现有方法多针对特定任务设计,且假设所有模态均完整可观测,但实际应用中常因各种因素导致模态缺失或不确定。传统模型过度关注缺失模态内部局部交互,忽视了多模态表示中的全局互补线索。为此,本文提出基于对比学习的动态多模态数据融合模型(CL-DMDF)。该模型引入一种新型跨特征与模态维度的注意力机制,有效计算可靠注意力得分,反映各层级的重要性。同时,设计实体-中心对比学习模块,利用实体特征构建基于中心的正样本,增强判别性学习。此外,采用自适应融合模块,提升动态融合策略的效率与准确性。在三个数据集上的大量实验表明,CL-DMDF在多样化的多模态融合任务中均表现优异。

原文摘要 · Abstract (English)

Multimodal data fusion involves integrating and analyzing information from multiple modalities to uncover latent correlations and complementary patterns, thereby enhancing data processing and decision-making. While existing methods for structured multimodal inputs are typically designed around specific tasks and assume fully observed modalities, real-world applications often suffer from uncertain or missing modality inputs due to various factors. Some traditional models overly emphasize local interactions within missing modalities, neglecting the global complementary cues embedded in multimodal representations. To overcome these limitations, we propose a Dynamic Multimodal Data Fusion model based on Contrastive Learning (CL-DMDF). CL-DMDF introduces a novel attention mechanism that operates across both feature and modality dimensions to compute reliable attention scores, effectively reflecting importance at each level. The CL-DMDF further incorporates an entity-centroid contrastive learning module that constructs centroid-based positive samples from entity features to enhance discriminative learning. Additionally, an adaptive fusion module is employed to improve the efficiency and accuracy of dynamic fusion strategies. Extensive experiments conducted on three datasets demonstrate the effectiveness of the CL-DMDF across diverse multimodal fusion tasks.

多模态融合对比学习动态融合缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。