首个音频地理定位基准数据集,测试大模型识图能力。
The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization
- 构建跨72国的音频定位数据集AGL1K,含1444个可定位音频片段。
- 闭源模型显著优于开源模型,语言线索主导定位判断。
- 适合研究音频-语言模型、地理推理与偏见分析的学者。
地理定位旨在推断信号的地理起源。在计算机视觉中,该任务是组合推理的严苛基准,且与公共安全密切相关。相比之下,音频地理定位因缺乏高质量音视频-位置配对数据而进展缓慢。为填补这一空白,我们提出AGL1K——首个面向音频语言模型(ALMs)的音频地理定位基准,覆盖72个国家和地区。为从众包平台中提取可靠可定位样本,我们设计了音频可定位性度量(Audio Localizability metric),量化每段录音的信息量,最终筛选出1,444个精调音频片段。在16个ALMs上评估发现,当前模型已具备音频地理定位能力。闭源模型显著优于开源模型,且语言线索常作为预测主干。我们进一步分析了模型推理路径、区域偏差、错误原因及度量可解释性。总体而言,AGL1K为音频地理定位建立基准,有望推动模型提升空间推理能力。
原文摘要 · Abstract (English)
Geo-localization aims to infer the geographic origin of a given signal. In computer vision, geo-localization has served as a demanding benchmark for compositional reasoning and is relevant to public safety. In contrast, progress on audio geo-localization has been constrained by the lack of high-quality audio-location pairs. To address this gap, we introduce AGL1K, the first audio geo-localization benchmark for audio language models (ALMs), spanning 72 countries and territories. To extract reliably localizable samples from a crowd-sourced platform, we propose the Audio Localizability metric that quantifies the informativeness of each recording, yielding 1,444 curated audio clips. Evaluations on 16 ALMs show that ALMs have emerged with audio geo-localization capability. We find that closed-source models substantially outperform open-source models, and that linguistic clues often dominate as a scaffold for prediction. We further analyze ALMs' reasoning traces, regional bias, error causes, and the interpretability of the localizability metric. Overall, AGL1K establishes a benchmark for audio geo-localization and may advance ALMs with better geospatial reasoning capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。