用机械臂动态调整麦克风阵列位置,提升嘈杂环境下的语音识别效果。
Lend me an Ear: Speech Enhancement Using a Robotic Arm with a Microphone Array
- 通过机械臂移动麦克风组,实时靠近说话人优化拾音
- 在不同信噪比下实现更高语音保真度和更低识别错误率
- 适合工业场景中对语音控制稳定性要求高的应用
噪声环境下语音增强性能显著下降,限制了语音控制技术在制造等工业场景的应用。现有方法主要依赖数字信号处理、深度学习或软件优化。本文提出一种新策略,通过物理层面动态调整麦克风阵列几何结构以适应声学变化。一个包含十六个麦克风的阵列安装于七自由度机械臂上,分为四组,其中一组靠近末端执行器。系统通过调节关节角度,使末端麦克风更接近目标说话人,从而提升参考信号质量。该方法融合声源定位、计算机视觉、逆运动学、最小方差无失真响应波束成形及深度神经网络时频掩码。实验表明,该方法优于传统固定配置,在多种输入信噪比条件下均实现更高的尺度不变信噪比与更低的词错误率。
原文摘要 · Abstract (English)
Speech enhancement performance degrades significantly in noisy environments, limiting the deployment of speech-controlled technologies in industrial settings, such as manufacturing plants. Existing speech enhancement solutions primarly rely on advanced digital signal processing techniques, deep learning methods, or complex software optimization techniques. This paper introduces a novel enhancement strategy that incorporates a physical optimization stage by dynamically modifying the geometry of a microphone array to adapt to changing acoustic conditions. A sixteen-microphone array is mounted on a robotic arm manipulator with seven degrees of freedom, with microphones divided into four groups of four, including one group positioned near the end-effector. The system reconfigures the array by adjusting the manipulator joint angles to place the end-effector microphones closer to the target speaker, thereby improving the reference signal quality. This proposed method integrates sound source localization techniques, computer vision, inverse kinematics, minimum variance distortionless response beamformer and time-frequency masking using a deep neural network. Experimental results demonstrate that this approach outperforms other traditional recording configruations, achieving higher scale-invariant signal-to-distortion ratio and lower word error rate accross multiple input signal-to-noise ratio conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。