用状态空间模型直接处理单比特语音信号,省去转换开销,适合低功耗设备。
Multirate State Space Models for End-to-End Processing of Pulse Density Modulated Speech Signals

- 利用状态空间模型的连续时间特性,直接处理脉冲密度调制信号。
- 在512kHz下实现稳定语音分类与增强,2MHz时性能媲美传统PCM方法。
- 输出可压缩超6.5万倍,大幅减少下游计算量,适合边缘设备部署。
基于状态空间模型(SSM)的深度神经网络在语音处理中日益普及,但通常作用于脉冲编码调制(PCM)音频。这限制了其在低功耗、始终在线的边缘设备上的应用,这些设备普遍采用单比特脉冲密度调制(PDM)微机电麦克风,因其噪声鲁棒性、低成本及可变采样率带来的低功耗优势。事实上,将PDM转换为PCM需低通滤波与降采样,给资源受限硬件带来高昂开销。尽管已有工作尝试直接处理PDM信号,但训练时间长且跨采样率泛化能力差。本文表明,SSM具备两项关键特性:其连续时间参数化可生成与调制方式和采样率无关的一致音频表示;其长期记忆使该表示可激进降采样而无需抗混叠操作。为此,我们提出一种新颖的端到端PDM语音处理架构,使用SSM将输入信号编码为调制与采样率无关的潜在表示。实验表明,该架构在512 kHz低功耗采样率下实现稳健的语音分类与增强,而在标准PDM采样率2 MHz下性能接近基于PCM的先进算法。此外,SSM输出可被压缩超过65,000倍,显著降低下游层的处理步数。
原文摘要 · Abstract (English)
Deep neural networks (DNNs) based on state-space models (SSMs) are increasingly applied to speech processing, but typically operate on pulse-code-modulated (PCM) audio. This constrains deployment on low-power, always-on edge devices, which commonly use single-bit pulse-density-modulated (PDM) micro-electromechanical (MEMS) microphones for their noise robustness, low cost, and variable sampling rates that enable low-power operation. In fact, converting PDM to PCM requires low-pass filtering and decimation, imposing costly overhead on resource-constrained hardware. While prior works have attempted to process PDM signals directly, they require long training times and generalize poorly across sampling rates. In this paper, we show that the SSM has two key properties that remediate these issues: its continuous-time parametrization allows it to produce a consistent representation of the input audio signal, regardless of the modulation strategy and sampling rate, and its long-term memory enables this representation to be aggressively downsampled without needing any anti-aliasing operations. We then propose a novel end-to-end PDM speech processing architecture that uses an SSM to encode the input audio signal into a modulation- and sampling-rate-invariant latent representation. We show that our proposed architecture achieves robust speech classification and enhancement gains at low-power sampling-rates (512 kHz) and similar performance to state-of-the-art algorithms operating on PCM data when tested on standard PDM sampling-rates of 2 MHz. Moreover, we show that the SSM's output can be downsampled by more than 65,000 times, thus significantly reducing the number of processing timesteps in downstream layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。