动态对齐语义的图信号框架,提升多模态融合鲁棒性
AGSP-DSA: An Adaptive Graph Signal Processing Framework for Robust Multimodal Fusion with Dynamic Semantic Alignment
- 构建双图结构捕捉模态内与跨模态关系,结合谱图滤波增强有效信号
- 在三个基准数据集上达到95.3%准确率,优于现有方法2.6个百分点
- 适用于缺失模态场景,适合情感分析与多媒体分类任务
本文提出自适应图信号处理与动态语义对齐框架(AGSP-DSA),用于异构多模态数据融合,包括文本、音频和图像。该方法采用双图构造学习模态内与跨模态关系,利用谱图滤波增强信息信号,并通过多尺度图卷积网络实现高效节点嵌入。引入语义感知注意力机制,使各模态可根据上下文相关性动态贡献。在CMU-MOSEI、AVE和MM-IMDB三个基准数据集上的实验表明,AGSP-DSA性能达状态领先水平:在CMU-MOSEI上取得95.3%准确率、0.936 F1-score和0.924 mAP,较MM-GNN提升2.6个百分点;在AVE上获93.4%准确率与0.911 F1-score;在MM-IMDB上达91.8%准确率与0.886 F1-score,验证了其在缺失模态设置下的良好泛化能力与鲁棒性,显著促进情感分析、事件识别与多媒体分类中的多模态学习。
原文摘要 · Abstract (English)
In this paper, we introduce an Adaptive Graph Signal Processing with Dynamic Semantic Alignment (AGSP DSA) framework to perform robust multimodal data fusion over heterogeneous sources, including text, audio, and images. The requested approach uses a dual-graph construction to learn both intra-modal and inter-modal relations, spectral graph filtering to boost the informative signals, and effective node embedding with Multi-scale Graph Convolutional Networks (GCNs). Semantic aware attention mechanism: each modality may dynamically contribute to the context with respect to contextual relevance. The experimental outcomes on three benchmark datasets, including CMU-MOSEI, AVE, and MM-IMDB, show that AGSP-DSA performs as the state of the art. More precisely, it achieves 95.3% accuracy, 0.936 F1-score, and 0.924 mAP on CMU-MOSEI, improving MM-GNN by 2.6 percent in accuracy. It gets 93.4% accuracy and 0.911 F1-score on AVE and 91.8% accuracy and 0.886 F1-score on MM-IMDB, which demonstrate good generalization and robustness in the missing modality setting. These findings verify the efficiency of AGSP-DSA in promoting multimodal learning in sentiment analysis, event recognition and multimedia classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。