通过置信度评估提升多模态情感识别的可靠性与鲁棒性。
A Trustworthy Method for Multimodal Emotion Recognition
- 基于模态置信度加权融合,提升决策可靠性。
- 在IEMOCAP和Music-video上分别达0.7511和0.9035的可信F1分数。
- 适用于噪声、异常或分布外数据场景,适合高可靠性需求应用。
现有情感识别方法多依赖复杂深度模型提升性能,但忽视了决策的可靠性,尤其在噪声、损坏及分布外数据下表现不佳。为此,我们提出可信情感识别(TER)方法,通过不确定性估计计算预测置信度,依据各模态置信度加权融合输出可信预测。同时提出新评估标准:引入可信精确率与召回率确定可信阈值,并构建可信准确率与可信F1分数评估模型可信性能。该框架结合置信度模块,赋予模型可靠性与抗噪鲁棒性。大量实验验证有效性:在Music-video数据集上达到82.40%准确率;在可信性能方面,于IEMOCAP和Music-video上分别取得0.7511与0.9035的可信F1分数,优于现有方法。
原文摘要 · Abstract (English)
Existing emotion recognition methods mainly focus on enhancing performance by employing complex deep models, typically resulting in significantly higher model complexity. Although effective, it is also crucial to ensure the reliability of the final decision, especially for noisy, corrupted and out-of-distribution data. To this end, we propose a novel emotion recognition method called trusted emotion recognition (TER), which utilizes uncertainty estimation to calculate the confidence value of predictions. TER combines the results from multiple modalities based on their confidence values to output the trusted predictions. We also provide a new evaluation criterion to assess the reliability of predictions. Specifically, we incorporate trusted precision and trusted recall to determine the trusted threshold and formulate the trusted Acc. and trusted F1 score to evaluate the model's trusted performance. The proposed framework combines the confidence module that accordingly endows the model with reliability and robustness against possible noise or corruption. The extensive experimental results validate the effectiveness of our proposed model. The TER achieves state-of-the-art performance on the Music-video, achieving 82.40% Acc. In terms of trusted performance, TER outperforms other methods on the IEMOCAP and Music-video, achieving trusted F1 scores of 0.7511 and 0.9035, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。