arXiv:2602.18527cs.CVcs.AI2026-02中稿 · ICML被引 2

让视觉语言模型学会在3D空间中定位声音,提升环境理解能力

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

  • 用RGB-D和多通道音频融合实现3D空间感知
  • 提出神经强度向量,在混响环境下仍能准确定位声源
  • 构建6.1万样本的基准数据集,适合研究物理环境推理

当前视听大模型主要基于2D感知,依赖RGB视频和单声道音频,导致在复杂3D环境中无法可靠进行声源定位与空间推理。为此,我们提出JAEGER框架,将视听大模型扩展至3D空间,通过融合RGB-D观测与多通道一阶全向音频,实现联合空间定位与推理。核心贡献是神经强度向量(Neural IV),一种可学习的空间音频表示,能在混叠声源等恶劣声学条件下仍提供稳健的方向线索,显著提升到达方向估计精度。为支持大规模训练与系统评估,我们构建了包含61,000条指令微调样本的SpatialSceneQA基准数据集,源自模拟物理环境。大量实验表明,本方法在多种空间感知与推理任务上均持续优于2D基准模型,证明显式3D建模对推动人工智能在物理环境中的发展至关重要。代码、预训练模型与数据集已开源。

原文摘要 · Abstract (English)

Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that precludes reliable source localization and spatial reasoning in complex 3D environments. We address this limitation by presenting JAEGER, a framework that extends AV-LLMs to 3D space, to enable joint spatial grounding and reasoning through the integration of RGB-D observations and multi-channel first-order ambisonics. A core contribution of our work is the neural intensity vector (Neural IV), a learned spatial audio representation that encodes robust directional cues to enhance direction-of-arrival estimation, even in adverse acoustic scenarios with overlapping sources. To facilitate large-scale training and systematic evaluation, we propose SpatialSceneQA, a benchmark of 61k instruction-tuning samples curated from simulated physical environments. Extensive experiments demonstrate that our approach consistently surpasses 2D-centric baselines across diverse spatial perception and reasoning tasks, underscoring the necessity of explicit 3D modelling for advancing AI in physical environments. Our source code, pre-trained model checkpoints, and datasets are available at https://github.com/liuzhan22/JAEGER.

3D感知音频定位多模态空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。