让视觉语言模型学会在动态环境中听声辨位,实现3D空间推理。
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
- 利用视听线索追踪目标物体轨迹,构建动态全局地图。
- 在真实场景中使现有模型的3D空间推理能力提升显著。
- 适合研究多模态感知、机器人导航与智能交互的学者。
动态视听环境中的三维空间推理是人类认知的核心,但现有视听大模型(AV-LLMs)和评测基准主要关注静态或二维场景。我们提出 SAVVY-Bench,首个针对动态场景中同步空间音频的3D空间推理评测基准。该基准包含数千个涉及静止与运动物体的关系,要求细粒度时间定位、一致的3D定位及多模态标注。为应对挑战,我们提出 SAVVY,一种无需训练的推理流程,包含两个阶段:(i) 第一人称空间轨迹估计,利用AV-LLMs及其他视听方法,结合视觉与空间音频线索追踪与查询相关的物体轨迹;(ii) 动态全局地图构建,聚合多模态查询对象轨迹并转化为统一的动态全局地图。通过坐标变换将全局地图与查询视角对齐,最终生成问答结果。实证评估表明,SAVVY显著提升现有顶尖AV-LLMs的性能,树立了动态3D空间推理的新标准。
原文摘要 · Abstract (English)
3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLMs) and benchmarks, which predominantly focus on static or 2D scenes. We introduce SAVVY-Bench, the first benchmark for 3D spatial reasoning in dynamic scenes with synchronized spatial audio. SAVVY-Bench is comprised of thousands of relationships involving static and moving objects, and requires fine-grained temporal grounding, consistent 3D localization, and multi-modal annotation. To tackle this challenge, we propose SAVVY, a novel training-free reasoning pipeline that consists of two stages: (i) Egocentric Spatial Tracks Estimation, which leverages AV-LLMs as well as other audio-visual methods to track the trajectories of key objects related to the query using both visual and spatial audio cues, and (ii) Dynamic Global Map Construction, which aggregates multi-modal queried object trajectories and converts them into a unified global dynamic map. Using the constructed map, a final QA answer is obtained through a coordinate transformation that aligns the global map with the queried viewpoint. Empirical evaluation demonstrates that SAVVY substantially enhances performance of state-of-the-art AV-LLMs, setting a new standard and stage for approaching dynamic 3D spatial reasoning in AV-LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。