多尺度多实例建模,精准定位视频中的发声物体
Multi-scale Multi-instance Visual Sound Localization and Segmentation
- 引入可学习的多尺度视觉特征,对齐音视频语义
- 在三个数据集上达到当前最优性能,显著提升定位精度
- 适合音视频分析、智能监控等需要声源定位的应用
视觉声音定位是预测视频中与声音源对应物体位置的典型难题。以往方法主要依赖全局音频与单尺度视觉特征之间的音视频关联来定位发声物体,尽管表现良好,但忽略了图像的多尺度视觉特征,且难以学习与真实标注具有区分性的区域。为此,本文提出一种新的多尺度多实例视觉声音定位框架M2VSL,能够直接从输入图像中学习与声音源相关的多尺度语义特征以定位发声物体。具体而言,M2VSL利用可学习的多尺度视觉特征,在图像多层级位置对齐音视频表示,并引入一种新型多尺度多实例注意力机制,动态聚合跨模态表示以实现定位。我们在VGGSound-Instruments、VGG-Sound Sources和AVSBench三个基准上进行了大量实验,结果表明所提M2VSL在发声物体定位与分割任务上均达到当前最优性能。
原文摘要 · Abstract (English)
Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global audio and one-scale visual features to localize sounding objects in each image. Despite their promising performance, they omitted multi-scale visual features of the corresponding image, and they cannot learn discriminative regions compared to ground truths. To address this issue, we propose a novel multi-scale multi-instance visual sound localization framework, namely M2VSL, that can directly learn multi-scale semantic features associated with sound sources from the input image to localize sounding objects. Specifically, our M2VSL leverages learnable multi-scale visual features to align audio-visual representations at multi-level locations of the corresponding image. We also introduce a novel multi-scale multi-instance transformer to dynamically aggregate multi-scale cross-modal representations for visual sound localization. We conduct extensive experiments on VGGSound-Instruments, VGG-Sound Sources, and AVSBench benchmarks. The results demonstrate that the proposed M2VSL can achieve state-of-the-art performance on sounding object localization and segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。