arXiv:2508.06902cs.CV2025-08被引 2

构建大规模短视频情感数据集并提出多模态融合网络,提升情感分析精度。

eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos

  • 提出多阶段标注流程与高质量数据集 eMotions,含27,996条视频。
  • 设计 AV-CANet 网络,融合音视频特征并引入局部-全局注意力机制。
  • 引入三极惩罚损失函数,有效缓解情感表达不一致问题,适合多模态研究者。

短视频已成为信息获取与分享的重要方式,其多模态复杂性对视频情感分析(VEA)提出新挑战。由于现有短视频情感数据稀缺,本文构建了 eMotions 大规模数据集,包含 27,996 条带全量标注的视频,通过多阶段标注流程保障质量、减少主观偏差,并提供类别均衡与面向测试的变体以满足多样化需求。针对短视频内容多样性带来的语义鸿沟及音视频共现不一致导致的局部偏差和信息缺失问题,提出端到端的音视频融合网络 AV-CANet,利用视频变换器捕捉语义相关表征,并设计局部-全局融合模块逐步建模跨模态关联。此外,构建 EP-CE 损失函数,通过三极惩罚全局引导优化。在三个 eMotions 相关数据集及四个公开 VEA 数据集上的实验验证了方法的有效性,同时通过消融实验分析关键组件。代码与数据将开源于 GitHub。

原文摘要 · Abstract (English)

Short-form videos (SVs) have become a vital part of our online routine for acquiring and sharing information. Their multimodal complexity poses new challenges for video analysis, highlighting the need for video emotion analysis (VEA) within the community. Given the limited availability of SVs emotion data, we introduce eMotions, a large-scale dataset consisting of 27,996 videos with full-scale annotations. To ensure quality and reduce subjective bias, we emphasize better personnel allocation and propose a multi-stage annotation procedure. Additionally, we provide the category-balanced and test-oriented variants through targeted sampling to meet diverse needs. While there have been significant studies on videos with clear emotional cues (e.g., facial expressions), analyzing emotions in SVs remains a challenging task. The challenge arises from the broader content diversity, which introduces more distinct semantic gaps and complicates the representations learning of emotion-related features. Furthermore, the prevalence of audio-visual co-expressions in SVs leads to the local biases and collective information gaps caused by the inconsistencies in emotional expressions. To tackle this, we propose AV-CANet, an end-to-end audio-visual fusion network that leverages video transformer to capture semantically relevant representations. We further introduce the Local-Global Fusion Module designed to progressively capture the correlations of audio-visual features. Besides, EP-CE Loss is constructed to globally steer optimizations with tripolar penalties. Extensive experiments across three eMotions-related datasets and four public VEA datasets demonstrate the effectiveness of our proposed AV-CANet, while providing broad insights for future research. Moreover, we conduct ablation studies to examine the critical components of our method. Dataset and code will be made available at Github.

情感分析多模态融合短视频数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。