arXiv:2603.12848cs.CVcs.AI2026-03被引 5

融合视觉、音频、文本多模态信息,提升无约束视频中犹豫情绪识别准确率。

Team LEYA in 10th ABAW Competition: Multimodal Ambivalence/Hesitancy Recognition Approach

  • 用四种模态互补建模:场景、人脸、音频、文本。
  • 多模态融合模型达到83.25%平均MF1,优于单一模态的70.02%。
  • 适合做情绪识别、人机交互等需要细粒度情感理解的研究者。

在无约束视频中识别犹豫/矛盾情绪是一项挑战性任务,因其行为特征微妙、多模态且高度依赖上下文。本文针对第10届ABAW竞赛,提出一种视频级犹豫/矛盾情绪识别的多模态方法。该方法融合四种互补模态:基于VideoMAE的场景动态、通过统计池化聚合的情绪帧嵌入的人脸信息、使用EmotionWav2Vec2.0提取并由Mamba模型处理的声学表示,以及微调的Transformer文本模型捕捉语言线索。各单模态嵌入再经多模态融合模型整合,包括原型增强变体。在BAH数据集上的实验表明,多模态融合显著优于所有单模态基线。最佳单模态配置达到70.02%平均MF1,最佳多模态融合模型达83.25%。最终测试最高性能为71.43%,由五个原型增强融合模型的集成获得。结果凸显互补多模态信号与鲁棒融合策略在识别犹豫情绪中的关键作用。

原文摘要 · Abstract (English)

Ambivalence/hesitancy recognition in unconstrained videos is a challenging problem due to the subtle, multimodal, and context-dependent nature of this behavioral state. In this paper, a multimodal approach for video-level ambivalence/hesitancy recognition is presented for the 10th ABAW Competition. The proposed approach integrates four complementary modalities: scene, face, audio, and text. Scene dynamics are captured with a VideoMAE-based model, facial information is encoded through emotional frame-level embeddings aggregated by statistical pooling, acoustic representations are extracted with EmotionWav2Vec2.0 and processed by a Mamba-based temporal encoder, and linguistic cues are modeled using fine-tuned transformer-based text models. The resulting unimodal embeddings are further combined using multimodal fusion models, including prototype-augmented variants. Experiments on the BAH corpus demonstrate clear gains of multimodal fusion over all unimodal baselines. The best unimodal configuration achieved an average MF1 of 70.02%, whereas the best multimodal fusion model reached 83.25%. The highest final test performance, 71.43%, was obtained by an ensemble of five prototype-augmented fusion models. The obtained results highlight the importance of complementary multimodal cues and robust fusion strategies for ambivalence/hesitancy recognition.

情绪识别多模态视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。