用耳机振动信号提升嘈杂环境语音质量,效果显著且适合移动端部署。
VibOmni: Towards Scalable Bone-conduction Speech Enhancement on Earables
- 融合麦克风音频与骨传导振动信号,实现多模态语音增强。
- 在真实数据集上提升21%语音质量得分,信噪比改善26%,误识率降40%。
- 仅需少量真实数据即可生成合成振动数据,适合资源受限设备使用。
真无线立体声耳机和VR/AR头戴设备日益普及,但其紧凑设计导致在嘈杂环境中进行通话或语音助手交互时性能受限。现有语音增强系统依赖全向麦克风,在背景噪声如多人对话中表现不佳。为此,我们提出VibOmni,一种轻量级端到端多模态语音增强系统,利用广泛可用的惯性测量单元(IMUs)采集骨传导振动信号。VibOmni采用双分支编码器-解码器网络融合音频与振动特征。为解决成对音-振数据稀缺问题,我们提出一种新型数据增强技术,基于有限录音建模骨传导函数(BCFs),仅需4.5%频谱相似性误差即可生成合成振动数据。此外,多模态信噪比(SNR)估计算法支持持续学习与自适应推理,无需设备端反向传播即可优化动态噪声环境下的性能。在32名志愿者使用不同设备的真实数据集上评估,VibOmni在感知语音质量(PESQ)上最高提升21%,信噪比(SNR)提升26%,字错误率(WER)降低约40%,且延迟更低。35人用户研究显示87%参与者更偏好VibOmni,证明其在多样化声学环境中的实用性。
原文摘要 · Abstract (English)
Earables, such as True Wireless Stereo earphones and VR/AR headsets, are increasingly popular, yet their compact design poses challenges for robust voice-related applications like telecommunication and voice assistant interactions in noisy environments. Existing speech enhancement systems, reliant solely on omnidirectional microphones, struggle with ambient noise like competing speakers. To address these issues, we propose VibOmni, a lightweight, end-to-end multi-modal speech enhancement system for earables that leverages bone-conducted vibrations captured by widely available Inertial Measurement Units (IMUs). VibOmni integrates a two-branch encoder-decoder deep neural network to fuse audio and vibration features. To overcome the scarcity of paired audio-vibration datasets, we introduce a novel data augmentation technique that models Bone Conduction Functions (BCFs) from limited recordings, enabling synthetic vibration data generation with only 4.5% spectrogram similarity error. Additionally, a multi-modal SNR estimator facilitates continual learning and adaptive inference, optimizing performance in dynamic, noisy settings without on-device back-propagation. Evaluated on real-world datasets from 32 volunteers with different devices, VibOmni achieves up to 21% improvement in Perceptual Evaluation of Speech Quality (PESQ), 26% in Signal-to-Noise Ratio (SNR), and about 40% WER reduction with much less latency on mobile devices. A user study with 35 participants showed 87% preferred VibOmni over baselines, demonstrating its effectiveness for deployment in diverse acoustic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。