arXiv:2501.08085cs.CLcs.LG2025-01被引 3

融合文本音频视频,用早期特征融合提升情感分类准确率

Dynamic Multimodal Sentiment Analysis: Leveraging Cross-Modal Attention for Enabled Classification

  • 在Transformer架构中尝试三种模态融合方式
  • 早期融合达71.87%准确率,优于后期融合
  • 多头注意力提升有限,适合研究模态交互的读者

本文研究一种融合文本、音频和视觉数据的多模态情感分析模型,旨在通过捕捉模态间的复杂交互关系,提升情感分类的准确性与细腻度。在基于Transformer的框架下,评估了三种特征融合策略:晚期融合、早期融合和多头注意力机制。实验基于包含同步文本、音频与视觉输入的CMU-MOSEI数据集,该数据集标注了情感得分。结果表明,早期融合显著优于晚期融合,准确率达到71.87%;多头注意力方法略有提升,达72.39%。研究发现,在当前框架下,早期模态融合有助于提升情感分类性能,而注意力机制的影响有限。未来工作将聚焦于优化特征融合技术、引入时间信息及探索动态特征加权以进一步提升模型表现。

原文摘要 · Abstract (English)

This paper explores the development of a multimodal sentiment analysis model that integrates text, audio, and visual data to enhance sentiment classification. The goal is to improve emotion detection by capturing the complex interactions between these modalities, thereby enabling more accurate and nuanced sentiment interpretation. The study evaluates three feature fusion strategies -- late stage fusion, early stage fusion, and multi-headed attention -- within a transformer-based architecture. Experiments were conducted using the CMU-MOSEI dataset, which includes synchronized text, audio, and visual inputs labeled with sentiment scores. Results show that early stage fusion significantly outperforms late stage fusion, achieving an accuracy of 71.87\%, while the multi-headed attention approach offers marginal improvement, reaching 72.39\%. The findings suggest that integrating modalities early in the process enhances sentiment classification, while attention mechanisms may have limited impact within the current framework. Future work will focus on refining feature fusion techniques, incorporating temporal data, and exploring dynamic feature weighting to further improve model performance.

多模态情感分析注意力机制特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。