arXiv:2506.03388cs.CV2025-06被引 1

对比街景与航拍影像,发现声音与视觉的匹配度差异。

Cross-Modal Urban Sensing: Evaluating Sound-Vision Alignment Across Street-Level and Aerial Imagery

  • 用音频-图像嵌入对比声景与视觉语义对应关系。
  • 街景嵌入比分割图更贴近环境声音,航拍分割更利于生态分类。
  • 为城市声景分析提供可解释的多模态方法,适合地理信息研究者。

环境声景蕴含丰富的城市生态与社会信息,但在大规模地理分析中尚未充分挖掘。本研究通过对比多种视觉表征策略在捕捉声学语义方面的表现,探究城市声音与视觉场景的对应程度。基于伦敦、纽约、东京三座全球主要城市,整合地理定位的音频记录与街景及遥感影像数据,采用AST模型处理音频,使用CLIP和RemoteCLIP提取图像特征,并结合CLIPSeg与Seg-Earth OV进行语义分割,以获取嵌入向量与类别级特征,评估跨模态相似性。结果表明,街景嵌入在声景语义对齐上优于分割输出,而遥感分割在基于生物声、地质声与人文声(BGA)框架下更有效识别生态类别。研究显示嵌入模型具有更强语义对齐能力,而分割方法则提供视觉结构与声景生态间的可解释关联。该工作推动了多模态城市感知的发展,为将声音纳入地理空间分析提供了新视角。

原文摘要 · Abstract (English)

Environmental soundscapes convey substantial ecological and social information regarding urban environments; however, their potential remains largely untapped in large-scale geographic analysis. In this study, we investigate the extent to which urban sounds correspond with visual scenes by comparing various visual representation strategies in capturing acoustic semantics. We employ a multimodal approach that integrates geo-referenced sound recordings with both street-level and remote sensing imagery across three major global cities: London, New York, and Tokyo. Utilizing the AST model for audio, along with CLIP and RemoteCLIP for imagery, as well as CLIPSeg and Seg-Earth OV for semantic segmentation, we extract embeddings and class-level features to evaluate cross-modal similarity. The results indicate that street view embeddings demonstrate stronger alignment with environmental sounds compared to segmentation outputs, whereas remote sensing segmentation is more effective in interpreting ecological categories through a Biophony--Geophony--Anthrophony (BGA) framework. These findings imply that embedding-based models offer superior semantic alignment, while segmentation-based methods provide interpretable links between visual structure and acoustic ecology. This work advances the burgeoning field of multimodal urban sensing by offering novel perspectives for incorporating sound into geospatial analysis.

多模态感知声景分析城市地理遥感影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。