arXiv:2604.14707cs.MMcs.SD2026-04

用卫星图生成真实地理音景,实现跨模态精准对齐。

Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery

论文配图:Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
图 1 · 摘自论文原文
  • 通过地理属性建模与声学假设扩展,统一生成音景。
  • 在20,000+样本上达到FAD 1.765,领先基线50%。
  • 适合城市规划、环境监测等需真实音景的应用。

近期图像到音频模型在以物体为中心的视觉场景中表现优异,但其在卫星影像上的应用受限于俯视视角复杂的语义模糊性。尽管卫星影像为全球音景生成提供了可扩展的数据源,但将其与具有独特空间结构的真实声学环境匹配仍具挑战。为此,我们提出Geo2Sound,一个新颖的任务与框架,用于从卫星影像生成地理真实的音景。该框架融合结构化地理属性建模、声学导向的语义假设扩展以及地理-声学对齐机制。轻量级分类器将高空场景归纳为紧凑地理属性,多组声学合理假设生成多样化候选音景,地理-声学对齐模块将地理属性投影至声学嵌入空间,识别最一致的候选。此外,我们构建了首个基准SatSound-Bench,包含超过20,000个高质量配对数据(卫星图像、文本描述、真实音频),覆盖10多个国家实地采集,并整合三个公开数据集。实验表明,Geo2Sound在FAD指标上达1.765,优于最强基线50.0%;人工评估显示真实感提升26.5%,语义一致性显著增强,验证了其大规模高保真合成能力。项目页面与源码:https://github.com/Blanketzzz/Geo2Sound

原文摘要 · Abstract (English)

Recent image-to-audio models have shown impressive performance on object-centric visual scenes. However, their application to satellite imagery remains limited by the complex, wide-area semantic ambiguity of top-down views. While satellite imagery provides a uniquely scalable source for global soundscape generation, matching these views to real acoustic environments with unique spatial structures is inherently difficult. To address this challenge, we introduce Geo2Sound, a novel task and framework for generating geographically realistic soundscapes from satellite imagery. Specifically, Geo2Sound combines structural geospatial attributes modeling, semantic hypothesis expansion, and geo-acoustic alignment in a unified framework. A lightweight classifier summarizes overhead scenes into compact geographic attributes, multiple sound-oriented semantic hypotheses are used to generate diverse acoustically plausible candidates, and a geo-acoustic alignment module projects geographic attributes into the acoustic embedding space and identifies the candidate most consistent with the candidate sets. Moreover, we establish SatSound-Bench, the first benchmark comprising over 20k high-quality paired satellite images, text descriptions, and real-world audio recordings, collected from the field across more than 10 countries and complemented by three public datasets. Experiments show that Geo2Sound achieves a SOTA FAD of 1.765, outperforming the strongest baseline by 50.0%. Human evaluations further confirm substantial gains in both realism (26.5%) and semantic alignment, validating our high-fidelity synthesis on scale. Project page and source code: https://github.com/Blanketzzz/Geo2Sound

音景生成卫星影像跨模态对齐地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。