arXiv:2503.00049cs.CVcs.AI2025-03被引 3

提出视频情感识别新任务,融合显隐场景信息精准定位情感来源。

Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos

  • 设计隐式增强因果混合专家模型,同时建模显性与隐性场景信息。
  • 在自建数据集上超越先进视频大模型,在情感定位精度上显著提升。
  • 适合关注视频情感分析、多模态理解的研究者和应用开发者。

视觉情感理解(VSU)研究长期依赖面部表情等显性场景信息判断情感,忽略了动作、物体关系和视觉背景等隐性信息,而这些对精准发现情感至关重要。为此,本文提出全新的视频全景场景驱动情感识别、定位与归因(Omni-SILA)任务,旨在通过显性与隐性场景信息的交互,精确识别、定位并归因视频中的情感。针对该任务的两大挑战——场景建模与隐性信息突出,本文提出隐式增强因果混合专家(ICM)方法。具体包括:场景平衡混合专家(SBM)模块用于建模场景,隐式增强因果(IEC)模块用于凸显隐性信息。在自建的显性与隐性场景数据集上,实验表明ICM方法显著优于先进视频大模型。

原文摘要 · Abstract (English)

Prior studies on Visual Sentiment Understanding (VSU) primarily rely on the explicit scene information (e.g., facial expression) to judge visual sentiments, which largely ignore implicit scene information (e.g., human action, objection relation and visual background), while such information is critical for precisely discovering visual sentiments. Motivated by this, this paper proposes a new Omni-scene driven visual Sentiment Identifying, Locating and Attributing in videos (Omni-SILA) task, aiming to interactively and precisely identify, locate and attribute visual sentiments through both explicit and implicit scene information. Furthermore, this paper believes that this Omni-SILA task faces two key challenges: modeling scene and highlighting implicit scene beyond explicit. To this end, this paper proposes an Implicit-enhanced Causal MoE (ICM) approach for addressing the Omni-SILA task. Specifically, a Scene-Balanced MoE (SBM) and an Implicit-Enhanced Causal (IEC) blocks are tailored to model scene information and highlight the implicit scene information beyond explicit, respectively. Extensive experimental results on our constructed explicit and implicit Omni-SILA datasets demonstrate the great advantage of the proposed ICM approach over advanced Video-LLMs.

情感识别视频理解多模态因果建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。