arXiv:2502.11478cs.SDcs.LG2025-02被引 6

发布首个喉部与麦克风配对语音数据集,助力降噪语音增强研究。

Throat and acoustic paired speech dataset for deep learning-based speech enhancement

  • 构建60名韩国母语者喉麦与普通麦克风同步录音的配对数据集
  • 提出最优对齐方法解决两类麦克风信号失配问题,提升语音质量
  • 验证基于映射的方法在恢复语音内容上更优,适合语音增强研究者使用

在工厂、地铁和繁忙街道等高噪声环境中,获取清晰语音极具挑战。喉部麦克风因其固有的降噪能力成为解决方案,但声波穿过皮肤和组织会衰减高频信息,降低语音清晰度。近年来深度学习方法在增强喉麦录音方面展现出潜力,但进展受限于缺乏标准数据集。本文介绍喉部与声学配对语音(TAPS)数据集,包含60名韩国母语者通过喉麦和声学麦克风录制的配对语音。此外,开发并应用了一种最优对齐方法,以解决两种麦克风间的固有信号失配问题。在TAPS数据集上测试了三种基线深度学习模型,发现基于映射的方法在提升语音质量和恢复语音内容方面表现更优。这些结果表明TAPS数据集在语音增强任务中的实用性,并支持其作为喉麦应用研究的标准资源。

原文摘要 · Abstract (English)

In high-noise environments such as factories, subways, and busy streets, capturing clear speech is challenging. Throat microphones can offer a solution because of their inherent noise-suppression capabilities; however, the passage of sound waves through skin and tissue attenuates high-frequency information, reducing speech clarity. Recent deep learning approaches have shown promise in enhancing throat microphone recordings, but further progress is constrained by the lack of a standard dataset. Here, we introduce the Throat and Acoustic Paired Speech (TAPS) dataset, a collection of paired utterances recorded from 60 native Korean speakers using throat and acoustic microphones. Furthermore, an optimal alignment approach was developed and applied to address the inherent signal mismatch between the two microphones. We tested three baseline deep learning models on the TAPS dataset and found mapping-based approaches to be superior for improving speech quality and restoring content. These findings demonstrate the TAPS dataset's utility for speech enhancement tasks and support its potential as a standard resource for advancing research in throat microphone-based applications.

语音增强数据集喉麦克风深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。