arXiv:2411.11454cs.CV2024-11被引 3

根据音视频语义相关性动态融合,提升视频显著性预测准确率

Relevance-guided Audio Visual Fusion for Video Saliency Prediction

  • 基于音视频语义相关性动态调整音频特征保留程度
  • 在6个数据集上显著优于现有方法,有效捕捉观众注意力分布
  • 适合关注视听融合、注意力建模的视觉智能研究者

音频与视频帧通常同步,对引导观众视觉注意力至关重要。将音频信息融入视频显著性预测可提升对人类视觉行为的预测能力。然而,现有方法常直接融合音视频特征,忽略了两者可能存在的不一致性(如背景音乐场景)。为此,本文提出新颖的关联引导音视频显著性预测网络AVRSP。具体而言,关联引导音视频特征融合模块(RAVF)根据音视频内容的语义相关性动态调节音频特征的保留程度,从而优化与视觉特征的融合过程。此外,多尺度特征协同模块(MS)整合不同编码阶段的视觉特征,增强对多尺度物体的表征能力;多尺度调节门(MRG)可将关键融合信息传递至视觉特征,优化多尺度视觉特征的利用效率。在六个音视频眼动数据集上的大量实验表明,所提AVRSP网络在音视频显著性预测任务中达到具有竞争力的性能。

原文摘要 · Abstract (English)

Audio data, often synchronized with video frames, plays a crucial role in guiding the audience's visual attention. Incorporating audio information into video saliency prediction tasks can enhance the prediction of human visual behavior. However, existing audio-visual saliency prediction methods often directly fuse audio and visual features, which ignore the possibility of inconsistency between the two modalities, such as when the audio serves as background music. To address this issue, we propose a novel relevance-guided audio-visual saliency prediction network dubbed AVRSP. Specifically, the Relevance-guided Audio-Visual feature Fusion module (RAVF) dynamically adjusts the retention of audio features based on the semantic relevance between audio and visual elements, thereby refining the integration process with visual features. Furthermore, the Multi-scale feature Synergy (MS) module integrates visual features from different encoding stages, enhancing the network's ability to represent objects at various scales. The Multi-scale Regulator Gate (MRG) could transfer crucial fusion information to visual features, thus optimizing the utilization of multi-scale visual features. Extensive experiments on six audio-visual eye movement datasets have demonstrated that our AVRSP network achieves competitive performance in audio-visual saliency prediction.

音视频融合显著性预测注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。