用扩散模型生成高质量多模态数据,提升情感分析鲁棒性。
QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis
- 通过扩散模型生成视觉和音频增强样本,扩大训练集。
- 在CH-SIMS上提升五分类准确率18.0%、二分类准确率5.9%。
- 自动加权低质量样本,适合数据稀缺场景下的多模态分析。
多模态大语言模型在捕捉多模态情感语义表示方面表现出强大能力,但其学习稳定且可泛化的多模态特征的能力受限于高质量训练数据的稀缺。为此,我们提出QASA(Quality-Aware Semantic Augmentation),利用扩散模型生成增强的视觉与听觉样本,扩充训练数据集,支持多模态学习。生成样本存在质量差异并可能出现跨模态不一致。为此,我们引入解耦的质量感知评分模块,根据每个增强样本的可靠性分配训练权重,降低低质量数据的影响,促进更稳定、鲁棒的模型训练。该框架结合扩散模型的生成能力与多模态大模型的语义推理能力,提供无需人工标注的自动化数据增强策略,在有限高质量数据条件下提升了模型泛化能力与鲁棒性。在CH-SIMS数据集上的实验显示,QASA在五分类准确率(Acc5)和二分类准确率(Acc2)上分别实现18.0%和5.9%的相对提升,并在CMU-MOSI与MUStARD基准上优于现有方法。
原文摘要 · Abstract (English)
Multimodal large language models have demonstrated strong ability in capturing semantic representations for multimodal sentiment analysis. Their capacity to learn stable and generalizable multimodal features is limited, however, by the scarcity of high-quality training data. To address this, we propose QASA (Quality-Aware Semantic Augmentation), which uses diffusion models to generate augmented visual and auditory samples, thereby enlarging the training dataset and supporting multimodal learning. The generated samples can vary in quality and may exhibit cross-modal inconsistencies. To manage this, we introduce a decoupled quality-aware scoring module that assigns training weights based on the reliability of each augmented sample. This approach reduces the influence of low-quality data and contributes to more stable and robust model training. The framework combines the generative capabilities of diffusion models with the semantic reasoning of multimodal large models, providing an automated data augmentation strategy that does not require human annotation while improving generalization and robustness under limited high-quality data. Experiments on the CH-SIMS dataset show that QASA yields a relative increase of 18.0\% and 5.9\% in five-class accuracy (Acc5) and binary accuracy (Acc2), respectively, and it also outperforms existing methods on the CMU-MOSI and MUStARD benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。