arXiv:2411.02236cs.CVcs.MM2024-11中稿 · NeurIPS被引 3

让声音物体在3D空间中精准分割,提升机器人与元宇宙的感知能力

3D Audio-Visual Segmentation

  • 将2D音视频分割扩展到3D空间,结合声场信息实现立体定位
  • 构建首个基于Habitat的3D音视频分割数据集,含34场景7类物体
  • 提出EchoSegnet模型,融合预训练视觉与3D场景表示,提升分割精度

在具身智能领域,识别场景中的发声物体是长期目标,广泛应用于机器人及AR/VR/MR。现有音视频分割(AVS)方法依赖同步摄像头与麦克风,从2D图像中定位发声物体掩码,但缺乏从2D到3D场景的映射,限制了实际应用。为此,本文提出全新的3D音视频分割任务,将输出拓展至3D空间。该任务面临相机外参变化、声音散射、遮挡及不同发声物体类别间声学差异等挑战。为推动研究,我们构建首个基于模拟器的基准数据集3DAVS-S34-O7,包含34个真实感3D场景与7类发声物体,支持单实例与多实例设置,通过重用Habitat模拟器生成物体位置与3D掩码标注。随后,提出EchoSegnet方法,通过空间音频感知的掩码对齐与优化,协同利用预训练2D音视频基础模型与3D视觉场景表示。大量实验表明,EchoSegnet在新基准上能有效实现3D空间中的发声物体分割,显著推进具身智能领域发展。

原文摘要 · Abstract (English)

Recognizing the sounding objects in scenes is a longstanding objective in embodied AI, with diverse applications in robotics and AR/VR/MR. To that end, Audio-Visual Segmentation (AVS), taking as condition an audio signal to identify the masks of the target sounding objects in an input image with synchronous camera and microphone sensors, has been recently advanced. However, this paradigm is still insufficient for real-world operation, as the mapping from 2D images to 3D scenes is missing. To address this fundamental limitation, we introduce a novel research problem, 3D Audio-Visual Segmentation, extending the existing AVS to the 3D output space. This problem poses more challenges due to variations in camera extrinsics, audio scattering, occlusions, and diverse acoustics across sounding object categories. To facilitate this research, we create the very first simulation based benchmark, 3DAVS-S34-O7, providing photorealistic 3D scene environments with grounded spatial audio under single-instance and multi-instance settings, across 34 scenes and 7 object categories. This is made possible by re-purposing the Habitat simulator to generate comprehensive annotations of sounding object locations and corresponding 3D masks. Subsequently, we propose a new approach, EchoSegnet, characterized by integrating the ready-to-use knowledge from pretrained 2D audio-visual foundation models synergistically with 3D visual scene representation through spatial audio-aware mask alignment and refinement. Extensive experiments demonstrate that EchoSegnet can effectively segment sounding objects in 3D space on our new benchmark, representing a significant advancement in the field of embodied AI. Project page: https://x-up-lab.github.io/research/3d-audio-visual-segmentation/

3D分割音视频融合具身智能仿真数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。