用卫星图生成声音地图,零样本实现全球声景建模
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
- 融合卫星图、文本描述和合成音频,跨模态学习声景概念
- 在GeoSound和SoundingEarth上超越现有方法,支持精准声景检索
- 可生成可播放的声景描述,适合教育与沉浸式应用
我们提出Sat2Sound,一个统一的多模态框架,用于地理空间声景理解,旨在预测并映射地球表面的声音分布。现有方法依赖配对的卫星图像与地理标记音频样本,难以涵盖位置的全部声音多样性。Sat2Sound通过引入由视觉-语言模型生成的语义丰富的声景描述来扩充数据集,扩展了每个位置可能代表的环境声音范围。该框架通过对比学习与码本对齐,联合学习音频、音频的文本描述、卫星图像及合成图像标题,发现跨模态共享的“声景概念”,实现高精度、可解释的局部声景制图。在GeoSound和SoundingEarth基准上,其跨模态检索性能达到当前最优。此外,通过检索可渲染为音频的详细声景描述,即使计算资源有限,也能实现基于位置的声景合成,适用于沉浸式与教育类应用。代码与模型已开源。
原文摘要 · Abstract (English)
We present Sat2Sound, a unified multimodal framework for geospatial soundscape understanding, designed to predict and map the distribution of sounds across the Earth's surface. Existing methods for this task rely on paired satellite images and geotagged audio samples, which often fail to capture the full diversity of sound at a location. Sat2Sound overcomes this limitation by augmenting datasets with semantically rich, vision-language model-generated soundscape descriptions, which broaden the range of possible ambient sounds represented at each location. Our framework jointly learns from audio, text descriptions of audio, satellite images, and synthetic image captions through contrastive and codebook-aligned learning, discovering a set of "soundscape concepts" shared across modalities, enabling hyper-localized, explainable soundscape mapping. Sat2Sound achieves state-of-the-art performance in cross-modal retrieval between satellite image and audio on the GeoSound and SoundingEarth benchmarks. Finally, by retrieving detailed soundscape captions that can be rendered through text-to-audio models, Sat2Sound enables location-conditioned soundscape synthesis for immersive and educational applications, even with limited computational resources. Our code and models are available at https://github.com/mvrl/sat2sound.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。