通过分布建模实现图像文本情感分析中的质量退化与缺失模态鲁棒处理。
Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and Fusion
- 基于特征分布队列统一建模低质与缺失模态
- 在三个数据集上优于当前最优方法,尤其在模态破坏场景下提升显著
- 适合实际社交平台中不完整或劣质图文数据的情感分析任务
随着社交媒体内容激增,对图文对中的情感进行分析已成为近年热点。尽管现有方法在融合图像与文本信息方面取得显著成果,但忽略了低质量及缺失模态的问题。在真实场景中,这类问题频繁发生,亟需具备鲁棒性的情感分析模型。为此,本文提出分布驱动的特征恢复与融合(DRF)方法,为图文对的鲁棒多模态情感分析提供支持。具体而言,为每种模态维护一个特征队列以逼近其特征分布,从而在统一框架下同时处理低质量和缺失模态。对于低质量模态,基于分布定量评估模态质量并降低其融合权重;对于缺失模态,利用样本与分布监督构建跨模态映射关系,从可用模态中恢复缺失信息。实验采用两种干扰策略,分别模拟不同现实场景下的模态损坏与丢弃。在三个公开图文数据集上的综合实验表明,相较于当前最优方法,DRF在两种策略下均实现普遍提升,验证了其在鲁棒多模态情感分析中的有效性。
原文摘要 · Abstract (English)
As posts on social media increase rapidly, analyzing the sentiments embedded in image-text pairs has become a popular research topic in recent years. Although existing works achieve impressive accomplishments in simultaneously harnessing image and text information, they lack the considerations of possible low-quality and missing modalities. In real-world applications, these issues might frequently occur, leading to urgent needs for models capable of predicting sentiment robustly. Therefore, we propose a Distribution-based feature Recovery and Fusion (DRF) method for robust multimodal sentiment analysis of image-text pairs. Specifically, we maintain a feature queue for each modality to approximate their feature distributions, through which we can simultaneously handle low-quality and missing modalities in a unified framework. For low-quality modalities, we reduce their contributions to the fusion by quantitatively estimating modality qualities based on the distributions. For missing modalities, we build inter-modal mapping relationships supervised by samples and distributions, thereby recovering the missing modalities from available ones. In experiments, two disruption strategies that corrupt and discard some modalities in samples are adopted to mimic the low-quality and missing modalities in various real-world scenarios. Through comprehensive experiments on three publicly available image-text datasets, we demonstrate the universal improvements of DRF compared to SOTA methods under both two strategies, validating its effectiveness in robust multimodal sentiment analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。