arXiv:2505.02331cs.CVcs.SD2025-05被引 11

用外部情感知识增强视觉音频情绪表征,提升识别精度。

VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection

  • 两阶段框架:先统一编码跨模态特征,再注入情感描述文本对齐语义。
  • 在多个数据集上达到顶尖性能,参数量小且泛化性强。
  • 适合需要高效、精准情绪识别的多模态应用开发者。

视听情绪识别(AVER)旨在从非语言的视听线索中推断人类情绪,具有模态互补和无需语言依赖的优势。然而,情绪表达的固有模糊性、跨模态表现差异以及高质量标注数据稀缺,使得该任务仍具挑战。现有自监督方法虽生成强多模态表征,但多依赖模态专用编码器和粗粒度内容对齐,难以建模细粒度情绪语义。为此,我们提出VAEmo,一种高效的两阶段情感中心联合视听表征学习框架,引入外部知识增强。第一阶段,在大规模以说话人为中心的视听语料上,通过掩码重建与对比学习预训练统一轻量级表征网络,缓解模态差距,学习互补表达且无需情绪标签。第二阶段,利用多模态大语言模型对少量样本设计思维链提示,生成详细情感描述;通过双路径对比学习将对应文本嵌入与视听表征对齐,进一步弥合情绪鸿沟。在多个下游AVER基准测试中,VAEmo实现最先进性能,结构紧凑,凸显统一跨模态编码与情感感知语义引导在高效、可泛化表征中的价值。

原文摘要 · Abstract (English)

Audiovisual emotion recognition (AVER) aims to infer human emotions from nonverbal visual-audio (VA) cues, offering modality-complementary and language-agnostic advantages. However, AVER remains challenging due to the inherent ambiguity of emotional expressions, cross-modal expressive disparities, and the scarcity of reliably annotated data. Recent self-supervised AVER approaches have introduced strong multimodal representations, yet they predominantly rely on modality-specific encoders and coarse content-level alignment, limiting fine-grained emotional semantic modeling. To address these issues, we propose VAEmo, an efficient two-stage framework for emotion-centric joint VA representation learning with external knowledge injection. In Stage~1, a unified and lightweight representation network is pre-trained on large-scale speaker-centric VA corpora via masked reconstruction and contrastive objectives, mitigating the modality gap and learning expressive, complementary representations without emotion labels. In Stage~2, multimodal large language models automatically generate detailed affective descriptions according to our well-designed chain-of-thought prompting for only a small subset of VA samples; these rich textual semantics are then injected by aligning their corresponding embeddings with VA representations through dual-path contrastive learning, further bridging the emotion gap. Extensive experiments on multiple downstream AVER benchmarks show that VAEmo achieves state-of-the-art performance with a compact design, highlighting the benefit of unified cross-modal encoding and emotion-aware semantic guidance for efficient, generalizable VA emotion representations.

情绪识别多模态知识注入自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。