解决多模态情感分析中缺失模态问题,提升实际应用泛化能力。
FSRF: Factorization-guided Semantic Recovery for Incomplete Multimodal Sentiment Analysis
- 通过因子分解分离同质、异质与噪声特征,增强表征学习。
- 双向知识迁移实现缺失语义的完整恢复,性能显著优于现有方法。
- 适合处理真实场景下模态不全的数据,尤其对隐私或设备限制场景有效。
近年来,多模态情感分析(MSA)成为热点,旨在利用多源数据理解人类情感。以往研究主要关注完整模态数据的交互与融合,忽视了真实应用中因遮挡、隐私限制或设备故障导致的模态缺失问题,造成模型泛化能力不足。为此,我们提出因子分解引导的语义恢复框架(FSRF),以缓解MSA任务中的模态缺失问题。具体地,设计去冗余的同质-异质因子分解模块,将模态分解为同质、异质及噪声表示,并构建精细约束机制用于表征学习;进一步设计分布对齐的自蒸馏模块,通过双向知识传递实现缺失语义的完整恢复。在两个数据集上的全面实验表明,当模态缺失情况不确定时,FSRF相比先前方法具有显著性能优势。
原文摘要 · Abstract (English)
In recent years, Multimodal Sentiment Analysis (MSA) has become a research hotspot that aims to utilize multimodal data for human sentiment understanding. Previous MSA studies have mainly focused on performing interaction and fusion on complete multimodal data, ignoring the problem of missing modalities in real-world applications due to occlusion, personal privacy constraints, and device malfunctions, resulting in low generalizability. To this end, we propose a Factorization-guided Semantic Recovery Framework (FSRF) to mitigate the modality missing problem in the MSA task. Specifically, we propose a de-redundant homo-heterogeneous factorization module that factorizes modality into modality-homogeneous, modality-heterogeneous, and noisy representations and design elaborate constraint paradigms for representation learning. Furthermore, we design a distribution-aligned self-distillation module that fully recovers the missing semantics by utilizing bidirectional knowledge transfer. Comprehensive experiments on two datasets indicate that FSRF has a significant performance advantage over previous methods with uncertain missing modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。