用文本做桥梁,让音频模型学会看图,无需成对数据。
Audio-to-Image Bird Species Retrieval without Audio-Image Pairs via Text Distillation
- 通过文本空间蒸馏,将图像模型的视觉语义迁移到音频模型。
- 在SSW60数据集上达到优于零样本组合基线的检索性能。
- 适合数据稀缺的生物声学物种识别场景,无需音频-图像配对数据。
音频到图像的检索为生物声学物种识别提供了可解释的替代方案,但受限于缺乏成对的音视频数据,学习对齐的音视频表示极具挑战。本文提出一种简单且数据高效的无监督方法,实现无需音频-图像配对数据的音频到图像检索。该方法利用文本作为语义中介:通过对比学习微调预训练音频-文本模型(BioLingual)的音频编码器,将其与预训练图像-文本模型(BioCLIP-2)的文本嵌入空间对齐,从而将富含视觉和分类结构的语义传递至音频表示中。这一蒸馏过程在训练中不使用图像,却能诱导音频与图像嵌入间的隐式对齐。我们在多个生物声学基准上评估模型表现,结果表明,蒸馏后的音频编码器在保留音频判别能力的同时,显著提升了焦点录音和声景数据集上的音频-文本对齐效果。最重要的是,在SSW60基准上,该方法实现了超越基于零样本模型组合或文本嵌入映射的基线的音频到图像检索性能,证明了通过文本进行间接语义传递足以实现有意义的音视频对齐,为数据稀缺环境下的视觉化物种识别提供了一种实用方案。
原文摘要 · Abstract (English)
Audio-to-image retrieval offers an interpretable alternative to audio-only classification for bioacoustic species recognition, but learning aligned audio-image representations is challenging due to the scarcity of paired audio-image data. We propose a simple and data-efficient approach that enables audio-to-image retrieval without any audio-image supervision. Our proposed method uses text as a semantic intermediary: we distill the text embedding space of a pretrained image-text model (BioCLIP-2), which encodes rich visual and taxonomic structure, into a pretrained audio-text model (BioLingual) by fine-tuning its audio encoder with a contrastive objective. This distillation transfers visually grounded semantics into the audio representation, inducing emergent alignment between audio and image embeddings without using images during training. We evaluate the resulting model on multiple bioacoustic benchmarks. The distilled audio encoder preserves audio discriminative power while substantially improving audio-text alignment on focal recordings and soundscape datasets. Most importantly, on the SSW60 benchmark, the proposed approach achieves strong audio-to-image retrieval performance exceeding baselines based on zero-shot model combinations or learned mappings between text embeddings, despite not training on paired audio-image data. These results demonstrate that indirect semantic transfer through text is sufficient to induce meaningful audio-image alignment, providing a practical solution for visually grounded species recognition in data-scarce bioacoustic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。