用超声波和视觉数据合成语音,无需人工标注即可提升语音清晰度。
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
- 通过音频桥接视频与超声波,建立跨模态对应关系。
- 合成超声数据在语音增强上媲美真实数据,性能超越现有方法。
- 适合做语音增强、人机交互的开发者和研究者使用。
语音增强对无处不在的人机交互至关重要。近年来,基于超声波的声学感知因其卓越的普适性和性能成为热门选择。然而,由于音频-超声数据采集过程中不可避免地存在意外干扰源,现有方法严重依赖人工进行数据采集与处理,导致数据稀缺,限制了超声语音增强的潜力。为此,我们提出 USpeech,一种无需人工干预的跨模态超声合成框架。其核心为两阶段架构:利用音频作为桥梁,建立视觉与超声模态间的对应关系,克服了缺乏配对视频-超声数据集及两者固有异构性的挑战。框架采用对比视频-音频预训练,将模态映射至共享语义空间,并使用音频-超声编码器-解码器生成超声数据。随后,构建时频域语音增强网络,并通过神经声码器恢复干净语音波形。大量实验表明,使用合成超声数据的 USpeech 性能可媲美真实数据,显著优于当前最先进基线方法。代码已开源:https://github.com/aiot-lab/USpeech/
原文摘要 · Abstract (English)
Speech enhancement is crucial for ubiquitous human-computer interaction. Recently, ultrasound-based acoustic sensing has emerged as an attractive choice for speech enhancement because of its superior ubiquity and performance. However, due to inevitable interference from unexpected and unintended sources during audio-ultrasound data acquisition, existing solutions rely heavily on human effort for data collection and processing. This leads to significant data scarcity that limits the full potential of ultrasound-based speech enhancement. To address this, we propose USpeech, a cross-modal ultrasound synthesis framework for speech enhancement with minimal human effort. At its core is a two-stage framework that establishes the correspondence between visual and ultrasonic modalities by leveraging audio as a bridge. This approach overcomes challenges from the lack of paired video-ultrasound datasets and the inherent heterogeneity between video and ultrasound data. Our framework incorporates contrastive video-audio pre-training to project modalities into a shared semantic space and employs an audio-ultrasound encoder-decoder for ultrasound synthesis. We then present a speech enhancement network that enhances speech in the time-frequency domain and recovers the clean speech waveform via a neural vocoder. Comprehensive experiments show USpeech achieves remarkable performance using synthetic ultrasound data comparable to physical data, outperforming state-of-the-art ultrasound-based speech enhancement baselines. USpeech is open-sourced at https://github.com/aiot-lab/USpeech/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。