不依赖标注数据,用预训练模型实现跨模态声音定位
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
- 用音频、视觉、文本三模态对齐,打通跨模态信息鸿沟
- 零样本下在DAVIS-AV、AVA-200等数据集上达到领先性能
- 适合无标注或标注稀缺的视频理解场景
音频视觉分割(AVS)旨在识别与声音源对应的视觉区域,在视频理解、监控和人机交互中具有重要作用。传统方法依赖大规模像素级标注,获取成本高且耗时。为此,我们提出一种新颖的零样本AVS框架,通过整合多个预训练模型,无需任务特定训练即可实现精准的声音源分割。该方法融合音频、视觉与文本表征,弥合模态间差异,实现无需AVS标注的细粒度跨模态分割。我们系统评估了不同预训练模型连接策略的有效性,并在多个数据集上验证其性能。实验结果表明,该框架在零样本条件下实现了当前最优的AVS表现,凸显了多模态模型集成在精细音频视觉分割中的有效性。
原文摘要 · Abstract (English)
Audiovisual segmentation (AVS) aims to identify visual regions corresponding to sound sources, playing a vital role in video understanding, surveillance, and human-computer interaction. Traditional AVS methods depend on large-scale pixel-level annotations, which are costly and time-consuming to obtain. To address this, we propose a novel zero-shot AVS framework that eliminates task-specific training by leveraging multiple pretrained models. Our approach integrates audio, vision, and text representations to bridge modality gaps, enabling precise sound source segmentation without AVS-specific annotations. We systematically explore different strategies for connecting pretrained models and evaluate their efficacy across multiple datasets. Experimental results demonstrate that our framework achieves state-of-the-art zero-shot AVS performance, highlighting the effectiveness of multimodal model integration for finegrained audiovisual segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。