用环境音生成真实地理场景,让声音能“画”出风景。
SounDiT: Geo-Contextual Soundscape-to-Landscape Generation
- 用扩散Transformer模型融合音景与地理信息生成景观
- 在两个大规模多模态数据集上实现高地理一致性生成
- 提出新评估指标PSS,从元素到感知全面衡量生成质量
近期的音频转图像模型在特定物体的图像生成方面表现优异,但无法根据环境音景重建真实世界景观。为填补这一空白,我们提出地理上下文音景到景观生成(GeoS2L)任务,旨在从环境音景合成具有地理真实性的景观图像。为此,我们构建了两个大规模地理上下文多模态数据集:SoundingSVI 和 SonicUrban,分别配对多样环境音景与真实景观图像。我们提出 SounDiT,一种基于扩散Transformer(DiT)的模型,结合环境音景与地理上下文场景条件,生成地理一致的景观图像。此外,我们提出位置相似度评分(Place Similarity Score, PSS),一个面向实际应用的地理上下文评估框架,用于衡量输入音景与生成图像之间的一致性。大量实验表明,SounDiT 在 GeoS2L 任务中优于现有基线,且 PSS 能有效捕捉元素、场景及人类感知层面的多层次生成一致性。
原文摘要 · Abstract (English)
Recent audio-to-image models have shown impressive performance in generating images of specific objects conditioned on their corresponding sounds. However, these models fail to reconstruct real-world landscapes conditioned on environmental soundscapes. To address this gap, we present Geo-contextual Soundscape-to-Landscape (GeoS2L) generation, a novel and practically significant task that aims to synthesize geographically realistic landscape images from environmental soundscapes. To support this task, we construct two large-scale geo-contextual multi-modal datasets, SoundingSVI and SonicUrban, which pair diverse environmental soundscapes with real-world landscape images. We propose SounDiT, a diffusion transformer (DiT)-based model that incorporates environmental soundscapes and geo-contextual scene conditioning to synthesize geographically coherent landscape images. Furthermore, we propose the Place Similarity Score (PSS), a practically-informed geo-contextual evaluation framework to measure consistency between input soundscapes and generated landscape images. Extensive experiments demonstrate that SounDiT outperforms existing baselines in the GeoS2L, while the PSS effectively captures multi-level generation consistency across element, scene,and human perception. Project page: https://gisense.github.io/SounDiT-Page/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。