让视觉音频语言联合理解场景的物体与位置关系
SceneBind: Binding What and Where Across Vision, Audio and Language

- 用语义-空间槽统一建模物体的'是什么'和'在哪儿'
- 跨模态检索与定位任务上达到新最好效果
- 适合需要精准空间感知的多模态应用
我们提出SceneBind,一种融合视觉、音频和语言的全景模态表征方法,实现对真实场景的语义与三维空间联合理解。现有全景编码器擅长识别物体(即‘是什么’),但缺乏显式空间结构(即‘在哪儿’)。SceneBind通过将每个场景表示为语义-空间实体,结合全局语义嵌入与以物体为中心的语义-空间槽,显式捕捉物体语义、空间属性及不确定性。我们进一步提出SceneBind Matching,一种融合全局场景相似性与物体对齐的匹配机制,支持跨模态场景检索与物体定位。为训练与评估,我们构建了一个新的真实世界双耳音频-视觉数据集,包含结构化语义与空间标注,并提出对齐多模态语义与空间信号的训练协议。SceneBind兼容大规模预训练语义编码器,仅需添加少量额外标记即可实现轻量级空间建模,在场景与空间检索任务上达到当前最优表现,并在音频-视觉定位等下游任务中展现出强大的零样本迁移能力。
原文摘要 · Abstract (English)
We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneBind Matching, a semantic-spatial matching scheme that integrates global scene similarity with object alignment, supporting cross-modal scene retrieval and object grounding. To train and evaluate SceneBind, we curate a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations, and propose a training protocol for aligning semantic and spatial signals across modalities. SceneBind is compatible with large-scale pretrained semantic encoders, adds lightweight spatial modeling with only a few additional tokens. It achieves state-of-the-art scene and spatial retrieval while enabling strong zero-shot transfer to downstream tasks such as audio-visual localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。