arXiv:2508.02974eess.AS2025-08被引 2

用神经音频编解码器提升喉麦语音清晰度,实时降噪效果佳。

Real-time speech enhancement in noise for throat microphone using neural audio codec as foundation model

  • 基于神经音频编解码器微调,实现喉麦语音实时增强。
  • 在Vibravox数据集上优于现有模型,显著提升语音可懂度。
  • 支持交互式界面,可实时切换增强、查看频谱与延迟。

我们展示了一个使用喉麦录制语音的实时语音增强演示系统。该系统完整呈现了从录音到深度学习后处理的全流程,适用于噪声环境下通过体传导麦克风采集的语音。喉麦捕捉皮肤振动,天然抑制外部噪声,但导致音频带宽受限。为此,我们在包含成对空气传导与喉麦录音的Vibravox数据集上,对支持实时推理的神经音频编解码器Kyutai的Mimi进行微调。相比当前最优模型,该策略展现出更优性能。推理过程集成于交互界面,用户可切换增强功能、可视化频谱图,并监控处理延迟。

原文摘要 · Abstract (English)

We present a real-time speech enhancement demo using speech captured with a throat microphone. This demo aims to showcase the complete pipeline, from recording to deep learning-based post-processing, for speech captured in noisy environments with a body-conducted microphone. The throat microphone records skin vibrations, which naturally attenuate external noise, but this robustness comes at the cost of reduced audio bandwidth. To address this challenge, we fine-tune Kyutai's Mimi--a neural audio codec supporting real-time inference--on Vibravox, a dataset containing paired air-conducted and throat microphone recordings. We compare this enhancement strategy against state-of-the-art models and demonstrate its superior performance. The inference runs in an interactive interface that allows users to toggle enhancement, visualize spectrograms, and monitor processing latency.

语音增强喉麦神经编解码器实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。