arXiv:2412.01488eess.AScs.LG2024-12被引 1

无需训练,用音频视觉联合分解实现声音驱动的精准图像分割

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization

  • 用非负矩阵分解联合分析预训练模型的音视频特征
  • 在未标注数据上达到当前最优的无监督分割效果
  • 适合追求高效部署与跨场景泛化的研究者

大规模预训练音视频模型展现出前所未有的泛化能力,适用于多种任务。本文针对声音驱动分割问题,即根据音频信号中听到的物体,定位图像中对应区域。现有方法多依赖微调或额外训练模块,而本文提出一种无需训练的方法:利用非负矩阵分解(NMF)对预训练模型提取的音视频特征进行联合因子分解,揭示共享的可解释语义概念,并将其输入开放词汇分割模型生成精确分割图。通过冻结预训练模型,本方法具备强泛化能力,在无监督声音驱动分割任务上达到当前最优性能,显著优于以往无监督方法。

原文摘要 · Abstract (English)

Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the specific task of sound-prompted segmentation, aiming to segment image regions corresponding to objects heard in an audio signal. Most existing approaches tackle this problem by fine-tuning pre-trained models or by training additional modules specifically for the task. We adopt a different strategy: we introduce a training-free approach that leverages Non-negative Matrix Factorization (NMF) to co-factorize audio and visual features from pre-trained models so as to reveal shared interpretable concepts. These concepts are passed on to an open-vocabulary segmentation model for precise segmentation maps. By using frozen pre-trained models, our method achieves high generalization and establishes state-of-the-art performance in unsupervised sound-prompted segmentation, significantly surpassing previous unsupervised methods.

声音分割音视频联合无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。