arXiv:2409.19595cs.SDcs.LG2024-09

提升音视频定位精度,强调声音特征重要性

Solution for Temporal Sound Localisation Task of ECCV Second Perception Test Challenge 2024

论文配图:Solution for Temporal Sound Localisation Task of ECCV Second Perception Test Challenge 2024
图 1 · 摘自论文原文
  • 优先强化声音模态特征,融合多模型提取音频信息
  • 最终测试得分0.4925,排名第一
  • 适合关注声音定位与多模态融合的研究者

本文提出一种改进的时序声音定位(Temporal Sound Localisation, TSL)方法,旨在根据预定义的声音类别对视频中的声音事件进行定位与分类。去年冠军方案采用音视频模态等权重融合,但本研究通过实验验证了声音特征在任务中的优势(第3节)。基于此,我们采用InterVideo、CaVMAE和VideoMAE等多种模型提取音频特征以增强声音模态表达。最终方法在决赛测试中取得0.4925的得分,位列第一。

原文摘要 · Abstract (English)

This report proposes an improved method for the Temporal Sound Localisation (TSL) task, which localizes and classifies the sound events occurring in the video according to a predefined set of sound classes. The champion solution from last year's first competition has explored the TSL by fusing audio and video modalities with the same weight. Considering the TSL task aims to localize sound events, we conduct relevant experiments that demonstrated the superiority of sound features (Section 3). Based on our findings, to enhance audio modality features, we employ various models to extract audio features, such as InterVideo, CaVMAE, and VideoMAE models. Our approach ranks first in the final test with a score of 0.4925.

声音定位多模态视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。