arXiv:2411.05603cs.CV2024-11被引 1

提出轻量级音视频融合方法,提升视频分类效率

Efficient Audio-Visual Fusion for Video Classification

  • 设计注意力机制实现高效音视频特征融合
  • 在YouTube-8M上性能接近大模型,参数量大幅减少
  • 适合资源受限场景的实时视频分类应用

我们提出Attend-Fusion,一种新颖且高效的视频分类任务中的音视频融合方法。该方法解决了如何有效利用音频和视觉模态的同时,保持紧凑的模型结构。在YouTube-8M数据集上的大量实验表明,Attend-Fusion在显著降低模型复杂度的前提下,实现了与更大规模基线模型相当的性能。

原文摘要 · Abstract (English)

We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visual modalities while maintaining a compact model architecture. Through extensive experiments on the YouTube-8M dataset, we demonstrate that our Attend-Fusion achieves competitive performance with significantly reduced model complexity compared to larger baseline models.

音视频融合视频分类轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。