arXiv:2507.21431cs.ROcs.HC2025-07被引 3

通过双麦克风系统实现户外人机交互中的精准语音定位。

Sound Source Localization for Human-Robot Interaction in Outdoor Environments

  • 结合粗对齐与时域回声消除,分离目标语音信号。
  • 在1dB信噪比下平均角度误差仅4度,95%精度达5度以内。
  • 适合嘈杂环境中需要定向交互的机器人应用。

本文提出一种基于无人地面车嵌入式麦克风阵列和操作员附近异步近讲麦克风的声音源定位策略。通过融合信号粗对齐与时域声学回声消除算法,估计出时频理想比率掩码,有效分离目标语音并抑制干扰与环境噪声。该方法支持选择性声音源定位,在信噪比为1dB时,平均角度误差为4度,95%定位精度达到5度以内,显著优于现有最先进方法。该技术可使机器人在嘈杂场景中准确获取操作者语音方向,实现丰富的人机交互。

原文摘要 · Abstract (English)

This paper presents a sound source localization strategy that relies on a microphone array embedded in an unmanned ground vehicle and an asynchronous close-talking microphone near the operator. A signal coarse alignment strategy is combined with a time-domain acoustic echo cancellation algorithm to estimate a time-frequency ideal ratio mask to isolate the target speech from interferences and environmental noise. This allows selective sound source localization, and provides the robot with the direction of arrival of sound from the active operator, which enables rich interaction in noisy scenarios. Results demonstrate an average angle error of 4 degrees and an accuracy within 5 degrees of 95\% at a signal-to-noise ratio of 1dB, which is significantly superior to the state-of-the-art localization methods.

语音定位人机交互麦克风阵列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。