arXiv:2409.09545cs.SDcs.LG2024-09中稿 · EUSIPCO 2025被引 2

多麦克风多模态融合提升混响环境情绪识别准确率

Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment

  • 用改进的音频-视频双流模型处理多通道声音与视频
  • 多麦克风多模态比单模态在混响下提升显著
  • 适用于语音情感分析、智能交互等实际场景

本文提出一种多模态情绪识别系统,旨在改善复杂声学条件下的情绪识别精度。方法结合改进扩展的分层令牌语义音频变换器(HTS-AT)处理多通道音频,以及R(2+1)D卷积神经网络(CNN)分析视频。在使用合成和真实房间冲击响应(RIRs)模拟混响的瑞尔森音频-视觉情绪语音与歌曲数据库(RAVDESS)上评估。结果表明,融合音频与视频模态的表现优于单模态方法,尤其在挑战性声学条件下;同时,采用多麦克风的多模态方法显著优于单麦克风版本。

原文摘要 · Abstract (English)

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio Transformer (HTS-AT) for multi-channel audio processing with an R(2+1)D Convolutional Neural Networks (CNN) model for video analysis. We evaluate our proposed method on a reverberated version of the Ryerson audio-visual database of emotional speech and song (RAVDESS) dataset using synthetic and real-world Room Impulse Responsess (RIRs). Our results demonstrate that integrating audio and video modalities yields superior performance compared to uni-modal approaches, especially in challenging acoustic conditions. Moreover, we show that the multimodal (audiovisual) approach that utilizes multiple microphones outperforms its single-microphone counterpart.

情绪识别多模态混响环境音频处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。