arXiv:2603.05708cs.CV2026-03被引 1

通过可解释的听觉感知实现高精度跨域定位

Interpretable Perception and Reasoning for Audiovisual Geolocation

  • 将音频分解为语义化的声学原子,提升听觉信号可解释性
  • 在20,000段视频上实现优于单模态基线的全球定位精度
  • 适合研究多模态定位、语音地理识别的科研人员

尽管多模态大语言模型(MLLM)在图像定位方面取得进展,但受视觉景观固有模糊性和听觉线索未被充分挖掘的影响,精确的全球地理定位仍具挑战。本文提出音频视觉定位框架,通过可解释的感知与推理解决地理模糊问题。构建了包含20,000个精心筛选片段、覆盖1,000个不同位置的大型视频基准集AVG。提出三阶段框架:(1) 使用混合自回归稀疏自编码器将噪声音频分解为语义化‘声学原子’;(2) 采用经过组相对策略优化(GRPO)微调的MLLM,融合声学原子与视觉特征进行多模态推理;(3) 在$S^2$流形上使用黎曼流匹配进行精度预测。实验表明,该框架显著优于单模态基线。结果表明,可解释的声景感知提供了关键且正交的信号,结合多模态推理可实现高精度全球定位。

原文摘要 · Abstract (English)

While recent advances in Multimodal Large Language Models (MLLMs) have improved image-based localization, precise global geolocation remains a formidable challenge due to the inherent ambiguity of visual landscapes and the largely untapped potential of auditory cues. In this paper, we introduce Audiovisual Geolocation, a framework designed to resolve geographic ambiguity through interpretable perception and reasoning. We present AVG, a high-quality global-scale video benchmark for geolocation, comprising 20,000 curated clips across 1,000 distinct locations. To address the complexity of audiovisual geolocation, we propose a three-stage framework: (1) a Perception stage that utilizes a mixture-autoregressive sparse autoencoder to decompose noisy audio into semantically grounded "acoustic atoms"; (2) a Multimodal Reasoning stage that employs an MLLM finetuned via Group Relative Policy Optimization (GRPO) to synthesize these atoms with visual features; and (3) a Precision Prediction stage using Riemannian Flow Matching on the $S^2$ manifold. Our experiments demonstrate that our framework significantly outperforms unimodal baselines. These results entail that interpretable perception of the soundscape provides a critical, orthogonal signal that, when coupled with multimodal reasoning, enables high-precision global localization.

多模态定位听觉感知地理识别可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。