用少麦克风实现高效3D声源定位,还能抗设备故障。
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
- 通过稀疏交叉注意力和预训练提升定位效率
- 仅需少量麦克风即可实现多声源定位
- 对麦克风位置误差有强容错能力,适合真实场景
声源定位(SSL)是确定复杂环境中声源位置的关键技术。现有方法存在计算开销大、需精确校准等问题,限制了其在动态或资源受限环境中的部署。本文提出一种新型3D SSL框架,采用稀疏交叉注意力、预训练及自适应信号相干性度量,在减少输入麦克风数量的前提下,实现高精度且计算高效的定位。该框架对不可靠甚至未知的麦克风位置输入具有故障容错能力,适用于真实场景。初步实验表明,其可扩展至多声源定位,无需额外硬件。本工作在模型性能与效率间取得平衡,并提升了真实场景下的鲁棒性。
原文摘要 · Abstract (English)
Sound source localization (SSL) is a critical technology for determining the position of sound sources in complex environments. However, existing methods face challenges such as high computational costs and precise calibration requirements, limiting their deployment in dynamic or resource-constrained environments. This paper introduces a novel 3D SSL framework, which uses sparse cross-attention, pretraining, and adaptive signal coherence metrics, to achieve accurate and computationally efficient localization with fewer input microphones. The framework is also fault-tolerant to unreliable or even unknown microphone position inputs, ensuring its applicability in real-world scenarios. Preliminary experiments demonstrate its scalability for multi-source localization without requiring additional hardware. This work advances SSL by balancing the model's performance and efficiency and improving its robustness for real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。