arXiv:2412.09317cs.SDcs.AI2024-12被引 4

融合音视频信息提升情感识别准确率,实验验证多模态方法有效性。

Multimodal Sentiment Analysis based on Video and Audio Inputs

  • 采用音视频双流模型分别处理音频与视频输入,融合概率进行决策。
  • 在CREMA-D和RAVDESS数据集上平均准确率达78.3%,关键模型表现突出。
  • 提出四种融合策略,适合做多模态情感分析的初学者参考。

尽管已有大量研究关注从视频和音频中进行情感分析,但找到最优模型以获得最高准确率仍是该领域的一大挑战。本文旨在验证同时使用视频和音频输入的情感识别模型的可行性。训练所用数据集为音频的CREMA-D和视频的RAVDESS。采用的预训练模型分别为:音频使用Facebook/wav2vec2-large,视频使用Google/vivit-b-16x2-kinetics400。通过取两个模型输出各情绪类别的概率平均值作为决策依据。当结果存在显著差异时,引入新测试框架,包括加权平均法、置信度阈值法、基于置信度的动态加权法和基于规则的逻辑法。该有限方法取得令人鼓舞的结果,表明未来对这些融合策略的研究具有可行性。

原文摘要 · Abstract (English)

Despite the abundance of current researches working on the sentiment analysis from videos and audios, finding the best model that gives the highest accuracy rate is still considered a challenge for researchers in this field. The main objective of this paper is to prove the usability of emotion recognition models that take video and audio inputs. The datasets used to train the models are the CREMA-D dataset for audio and the RAVDESS dataset for video. The fine-tuned models that been used are: Facebook/wav2vec2-large for audio and the Google/vivit-b-16x2-kinetics400 for video. The avarage of the probabilities for each emotion generated by the two previous models is utilized in the decision making framework. After disparity in the results, if one of the models gets much higher accuracy, another test framework is created. The methods used are the Weighted Average method, the Confidence Level Threshold method, the Dynamic Weighting Based on Confidence method, and the Rule-Based Logic method. This limited approach gives encouraging results that make future research into these methods viable.

情感分析多模态音视频融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。