用地理语义信息提升音频事件识别准确率,解决声音混淆难题。
Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context
- 结合地理语义上下文与音频特征进行多标签声景识别
- 在28类事件上实现10.71小时标注音频的跨模态融合评测
- 模型表现接近人类水平,适合研究地理+音频联合建模
计算听觉场景分析(CASA)中的环境声音理解通常被当作纯音频识别问题处理,导致多标签音频标记(AT)中因声学相似性难以区分某些事件。此时,消歧线索往往存在于波形之外。基于地理信息系统数据(如兴趣点)生成的地理语义上下文(GSC)可提供与位置相关的环境先验,有助于降低此类模糊性。为此,本文提出地理音频标记(Geo-AT)任务,将多标签声景标记同时依赖于音频和GSC。为评测该任务,构建了包含10.71小时音频、28个事件类别、每段音频配有11类语义上下文表示的基准数据集Geo-ATBench。提出统一的地理-音频融合框架GeoFusion-AT,评估特征级、表示级和决策级融合策略,并对比音频与GSC单模态基线。结果表明,引入GSC显著提升标记性能,尤其在声学混淆标签上;10名参与者对579个样本的众包听觉研究表明,模型在Geo-ATBench上的表现与人工标注聚合结果无显著差异,验证其人类对齐性。该任务、基准与可复现框架为CASA社区研究地理语义上下文下的音频标记奠定了基础。数据集、代码与模型已开源。
原文摘要 · Abstract (English)
Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent drawback in multi-label audio tagging (AT): acoustic similarity can make certain events difficult to separate from waveforms alone. In such cases, disambiguating cues often lie outside the waveform. Geospatial semantic context (GSC), derived from geographic information system data, e.g., points of interest (POI), provides location-tied environmental priors that can help reduce this ambiguity. A systematic study of this direction is enabled through the proposed geospatial audio tagging (Geo-AT) task, which conditions multi-label sound event tagging on GSC alongside audio. To benchmark Geo-AT, Geo-ATBench is introduced as a polyphonic audio benchmark with geographical annotations, containing 10.71 hours of audio across 28 event categories; each clip is paired with a GSC representation from 11 semantic context categories. GeoFusion-AT is proposed as a unified geo-audio fusion framework that evaluates feature-, representation-, and decision-level fusion on representative audio backbones, with audio- and GSC-only baselines. Results show that incorporating GSC improves AT performance, especially on acoustically confounded labels, indicating geospatial semantics provide effective priors beyond audio alone. A crowdsourced listening study with 10 participants on 579 samples shows that there is no significant difference in performance between models on Geo-ATBench labels and aggregated human labels, supporting Geo-ATBench as a human-aligned benchmark. The Geo-AT task, benchmark Geo-ATBench, and reproducible geo-audio fusion framework GeoFusion-AT provide a foundation for studying AT with geospatial semantic context within the CASA community. Dataset, code, models are on homepage (https://github.com/WuYanru2002/Geo-ATBench).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。