arXiv:2409.06274cs.ROcs.SD2024-09被引 1

用双掩码生成对抗网络修复机器人自语过滤中的语音失真

Spectral oversubtraction? An approach for speech enhancement after robot ego speech filtering in semi-real-time

  • 设计双掩码Conformer结构,融合高频信息与长期特征补偿基频区失真
  • 在两个数据集上实现识别准确率显著提升,含未见噪声场景表现优异
  • 支持流式输入的增量处理,适用于实时人机交互场景

谱减法因简单性被广泛用于解决机器人自语过滤(RESF)问题,即从单麦克风录音中检测人类打断语音。然而该方法在基频区(FFR)存在过减现象,导致语音识别性能下降。为此,本文提出基于双掩码Conformer的度量生成对抗网络(CMGAN),利用高频信息与长期特征补偿被过度抑制的基频区,并对重构谱图进行去噪。同时引入增量处理机制,使在固定长输入训练的模型可支持流式音频的半实时处理。在两个数据集上的评估表明,所提方法在识别准确率上显著提升,双掩码策略与增量处理均有效增强了实际人机交互场景下RESF系统的鲁棒性。

原文摘要 · Abstract (English)

Spectral subtraction, widely used for its simplicity, has been employed to address the Robot Ego Speech Filtering (RESF) problem for detecting speech contents of human interruption from robot's single-channel microphone recordings when it is speaking. However, this approach suffers from oversubtraction in the fundamental frequency range (FFR), leading to degraded speech content recognition. To address this, we propose a Two-Mask Conformer-based Metric Generative Adversarial Network (CMGAN) to enhance the detected speech and improve recognition results. Our model compensates for oversubtracted FFR values with high-frequency information and long-term features and then de-noises the new spectrogram. In addition, we introduce an incremental processing method that allows semi-real-time audio processing with streaming input on a network trained on long fixed-length input. Evaluations of two datasets, including one with unseen noise, demonstrate significant improvements in recognition accuracy and the effectiveness of the proposed two-mask approach and incremental processing, enhancing the robustness of the proposed RESF pipeline in real-world HRI scenarios.

语音增强生成对抗网络人机交互实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。