用多视角音视频设备定位不可见声源并分类,提升工业检测精度。
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
- 将声源定位建模为集合预测问题,融合单视角音频与多视角视觉信息。
- 在模拟数据集上实现高精度3D定位,对噪声和传感器误差鲁棒。
- 适合做设备故障检测、气体泄漏等需要听觉感知的场景应用。
准确地定位3D声源并估计其语义标签——即使声源不可见,但假设其位于场景中物体的物理表面——在气体泄漏检测、机械故障诊断等实际应用中具有重要意义。由于音视频之间存在弱相关性,如何利用跨模态信息成为新挑战。为此,我们提出使用包含针孔RGB-D相机与共面四麦克风阵列(Mic-Array)的声学相机装置,从多视角同步采集音视频信号,通过跨模态线索估计声源3D位置。具体而言,我们的框架SoundLoc3D将任务视为集合预测问题,集合中的每个元素对应一个潜在声源。初始时,集合表示由单视角麦克风信号学习得到,随后通过主动融合多视角RGB-D图像揭示的物理表面约束进行优化。我们在大规模模拟数据集上验证了SoundLoc3D的有效性与优越性,并进一步展示了其对RGB-D测量误差和环境噪声干扰的鲁棒性。
原文摘要 · Abstract (English)
Accurately localizing 3D sound sources and estimating their semantic labels -- where the sources may not be visible, but are assumed to lie on the physical surface of objects in the scene -- have many real applications, including detecting gas leak and machinery malfunction. The audio-visual weak-correlation in such setting poses new challenges in deriving innovative methods to answer if or how we can use cross-modal information to solve the task. Towards this end, we propose to use an acoustic-camera rig consisting of a pinhole RGB-D camera and a coplanar four-channel microphone array~(Mic-Array). By using this rig to record audio-visual signals from multiviews, we can use the cross-modal cues to estimate the sound sources 3D locations. Specifically, our framework SoundLoc3D treats the task as a set prediction problem, each element in the set corresponds to a potential sound source. Given the audio-visual weak-correlation, the set representation is initially learned from a single view microphone array signal, and then refined by actively incorporating physical surface cues revealed from multiview RGB-D images. We demonstrate the efficiency and superiority of SoundLoc3D on large-scale simulated dataset, and further show its robustness to RGB-D measurement inaccuracy and ambient noise interference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。