arXiv:2501.09502cs.CV2025-01被引 39

提升视频多模态模型情感分析能力,融合面部微表情与音频细节

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

  • 将面部编码模型嵌入视频多模态大模型,统一处理音视频情绪线索
  • 在24,137个粗粒度样本和3,500个细粒度标注数据上实现最优性能
  • 适合需要高精度情感理解的交互系统、智能客服等应用场景

准确理解情绪对人机交互等领域至关重要。由于情绪具有多模态特性(如受面部表情和音频影响),研究者转向使用多模态模型而非单模态方法。然而,现有视频多模态大语言模型(Video MLLM)在融合音频与识别细微面部微表情方面仍存在困难,且缺乏详细的情绪分析数据集限制了发展。为此,我们构建了自审数据集和人工审核数据集,分别包含24,137个粗粒度样本和3,500个手动标注的细粒度情绪样本,支持模型学习多样化场景并更好泛化至真实应用。此外,除音频建模外,我们提出显式集成面部编码模型到先进Video MLLM中,使模型能有效统一音频与细微面部线索进行情绪理解。通过在统一空间对齐特征并采用指令微调,Omni-Emotion在情绪识别与推理任务中均达到当前最佳表现。

原文摘要 · Abstract (English)

Understanding emotions accurately is essential for fields like human-computer interaction. Due to the complexity of emotions and their multi-modal nature (e.g., emotions are influenced by facial expressions and audio), researchers have turned to using multi-modal models to understand human emotions rather than single-modality. However, current video multi-modal large language models (MLLMs) encounter difficulties in effectively integrating audio and identifying subtle facial micro-expressions. Furthermore, the lack of detailed emotion analysis datasets also limits the development of multimodal emotion analysis. To address these issues, we introduce a self-reviewed dataset and a human-reviewed dataset, comprising 24,137 coarse-grained samples and 3,500 manually annotated samples with detailed emotion annotations, respectively. These datasets allow models to learn from diverse scenarios and better generalize to real-world applications. Moreover, in addition to the audio modeling, we propose to explicitly integrate facial encoding models into the existing advanced Video MLLM, enabling the MLLM to effectively unify audio and the subtle facial cues for emotion understanding. By aligning these features within a unified space and employing instruction tuning in our proposed datasets, our Omni-Emotion achieves state-of-the-art performance in both emotion recognition and reasoning tasks.

情感分析多模态模型视频理解面部编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。