低复杂度语音重建让耳机在嘈杂环境中清晰捕捉用户自声。
Low-Complexity Own Voice Reconstruction for Hearables with an In-Ear Microphone
- 基于FT-JNF架构设计轻量级语音重建模型,适配耳机算力限制。
- 仅需少量设备专属录音即可实现高质量语音重建,信噪比提升显著。
- 适合资源受限的可穿戴设备,尤其适用于嘈杂环境下的语音通信。
可穿戴设备通常配备一个或多个麦克风,用于语音通信。本文聚焦于耳机在嘈杂环境中捕捉用户自声的应用场景。在此场景下,自声重建(OVR)对提升录制语音的质量和可懂度至关重要。此前工作已开发基于深度学习的OVR系统,通过语音音素依赖的自声传输特性建模进行数据增强,以减少设备专属录音的数据需求。鉴于可穿戴设备计算资源有限,本文提出基于FT-JNF架构的低复杂度OVR变体,并研究有效数据增强与微调所需的设备专属录音量。仿真结果表明,所提系统在低复杂度和少量设备专属数据条件下仍能显著提升语音质量。
原文摘要 · Abstract (English)
Hearable devices, equipped with one or more microphones, are commonly used for speech communication. Here, we consider the scenario where a hearable is used to capture the user's own voice in a noisy environment. In this scenario, own voice reconstruction (OVR) is essential for enhancing the quality and intelligibility of the recorded noisy own voice signals. In previous work, we developed a deep learning-based OVR system, aiming to reduce the amount of device-specific recordings for training by using data augmentation with phoneme-dependent models of own voice transfer characteristics. Given the limited computational resources available on hearables, in this paper we propose low-complexity variants of an OVR system based on the FT-JNF architecture and investigate the required amount of device-specific recordings for effective data augmentation and fine-tuning. Simulation results show that the proposed OVR system considerably improves speech quality, even under constraints of low complexity and a limited amount of device-specific recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。