arXiv:2504.03373cs.SDcs.RO2025-04

用GPU加速噪声鲁棒声源定位,让机器人听觉实时运行。

An Efficient GPU-based Implementation for Noise Robust Sound Source Localization

  • 在HARK平台用GPU实现基于GSVD-MUSIC的声源定位
  • 60麦克风阵列下,嵌入式设备提速5648倍,全模块快10.7倍
  • 适合需实时处理大阵列音频的机器人、智能设备开发

机器人听觉包括声源定位(SSL)、声源分离(SSS)和自动语音识别(ASR),使机器人和智能设备具备类人听觉能力。尽管应用广泛,但多通道麦克风阵列的SSL处理涉及大量计算密集型矩阵运算,在资源受限的中央处理器(CPU)上部署效率低。本文提出在开源软件平台HARK中,基于通用奇异值分解的多重信号分类(GSVD-MUSIC)算法,实现面向机器人听觉的GPU加速版SSL。针对60通道麦克风阵列,实验表明:在搭载NVIDIA GPU与ARM Cortex-A78AE v8.2 64位CPU的Jetson AGX Orin嵌入式设备上,GSVD计算速度提升5648.7倍,整个SSL模块提速10.7倍;在配备NVIDIA A100 GPU与AMD EPYC 7352 CPU的服务器上,对应速度提升分别为4245.1倍和17.3倍,使大规模麦克风阵列的实时处理成为可能,并为后续机器学习或深度学习任务预留充足计算余量。

原文摘要 · Abstract (English)

Robot audition, encompassing Sound Source Localization (SSL), Sound Source Separation (SSS), and Automatic Speech Recognition (ASR), enables robots and smart devices to acquire auditory capabilities similar to human hearing. Despite their wide applicability, processing multi-channel audio signals from microphone arrays in SSL involves computationally intensive matrix operations, which can hinder efficient deployment on Central Processing Units (CPUs), particularly in embedded systems with limited CPU resources. This paper introduces a GPU-based implementation of SSL for robot audition, utilizing the Generalized Singular Value Decomposition-based Multiple Signal Classification (GSVD-MUSIC), a noise-robust algorithm, within the HARK platform, an open-source software suite. For a 60-channel microphone array, the proposed implementation achieves significant performance improvements. On the Jetson AGX Orin, an embedded device powered by an NVIDIA GPU and ARM Cortex-A78AE v8.2 64-bit CPUs, we observe speedups of 5648.7x for GSVD calculations and 10.7x for the SSL module, while speedups of 4245.1x for GSVD calculation and 17.3x for the entire SSL module on a server configured with an NVIDIA A100 GPU and AMD EPYC 7352 CPUs, making real-time processing feasible for large-scale microphone arrays and providing ample capacity for real-time processing of potential subsequent machine learning or deep learning tasks.

声源定位GPU加速机器人听觉实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。