arXiv:2410.02804cs.CVcs.AI2024-10被引 3

用相似数据检索填补缺失模态,提升情绪识别准确率

Leveraging Retrieval Augment Approach for Multimodal Emotion Recognition Under Missing Modalities

  • 引入外部数据库检索相似多模态数据补全缺失信息
  • 在多个数据集上显著优于现有最先进方法
  • 适合处理传感器故障导致模态缺失的实际场景

多模态情绪识别依赖完整的多模态信息和鲁棒的联合表示以获得高性能。然而现实中常出现模态缺失,如因传感器故障或网络带宽问题导致视频、音频或文本数据缺失,给情绪识别研究带来挑战。传统方法通过完整模态提取信息并重建缺失模态来学习联合表示,虽有一定成效,但仅依赖内部重构与联合学习,在关键信息缺失时仍存在局限。为此,我们提出检索增强的缺失模态多模态情绪识别框架(RAMER),通过引入包含相关多模态情绪数据的数据库,检索相似样本以填补缺失模态的信息空缺。大量实验表明,该框架在缺失模态情绪识别任务中优于现有最先进方法。项目代码已开源:https://github.com/WooyoohL/Retrieval_Augment_MER。

原文摘要 · Abstract (English)

Multimodal emotion recognition utilizes complete multimodal information and robust multimodal joint representation to gain high performance. However, the ideal condition of full modality integrity is often not applicable in reality and there always appears the situation that some modalities are missing. For example, video, audio, or text data is missing due to sensor failure or network bandwidth problems, which presents a great challenge to MER research. Traditional methods extract useful information from the complete modalities and reconstruct the missing modalities to learn robust multimodal joint representation. These methods have laid a solid foundation for research in this field, and to a certain extent, alleviated the difficulty of multimodal emotion recognition under missing modalities. However, relying solely on internal reconstruction and multimodal joint learning has its limitations, especially when the missing information is critical for emotion recognition. To address this challenge, we propose a novel framework of Retrieval Augment for Missing Modality Multimodal Emotion Recognition (RAMER), which introduces similar multimodal emotion data to enhance the performance of emotion recognition under missing modalities. By leveraging databases, that contain related multimodal emotion data, we can retrieve similar multimodal emotion information to fill in the gaps left by missing modalities. Various experimental results demonstrate that our framework is superior to existing state-of-the-art approaches in missing modality MER tasks. Our whole project is publicly available on https://github.com/WooyoohL/Retrieval_Augment_MER.

情绪识别多模态缺失模态检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。