通过亚奈奎斯特采样与宽带重建,大幅降低骨传导耳机功耗。
CAPS: A Cascaded Reconstruction Model to Power Saving in Hearables Using Sub-Nyquist Sampling with Bandwidth Extension

- 采用低采样率和低比特分辨率的模数转换器,实现功耗下降3.3倍。
- 在窄带信号下重建宽带语音,保障真实场景下语音可懂度。
- 推理延迟仅1.36毫秒,内存占用11.04MB,适合移动端部署。
耳戴设备是佩戴于耳朵上的可穿戴计算机。在嘈杂环境下,耳戴设备通常结合骨传导麦克风与空气传导麦克风进行多模态语音增强。然而,现有模型大多未探究在模数转换器(ADC)中同时降低采样位分辨率与采样频率对功耗与音质的影响。此外,当前框架无法在耳戴设备上实现亚奈奎斯特采样,因缺乏从窄带分量重建宽带信号的方法。为此,我们提出CAPS,(i)主动采用亚奈奎斯特采样与低比特分辨率的ADC,使耳戴设备功耗降低3.3倍;(ii)支持移动平台流式运行,推理时间仅为1.36毫秒,内存占用11.04MB。CAPS在真实环境中确保了语音可懂度,弥合了效率与节能之间的差距。
原文摘要 · Abstract (English)
Hearables are wearable computers worn on the ear. Bone conduction microphones are used with air conduction microphones in hearables for multimodal speech enhancement in noisy conditions. Despite this potential, current models largely fail to explore how jointly reducing sampling bit resolution and sampling frequency in analog-to-digital converters (ADCs) of hearables impacts both power usage and audio quality. Furthermore, current frameworks cannot do sub-Nyquist sampling in hearables because they lack a method to reconstruct wideband signals from narrowband components. We therefore propose CAPS, which (i) intentionally employs sub-Nyquist sampling and low bit resolution in ADCs, achieving a 3.3x reduction in power consumption in hearables, and (ii) supports streaming operation on mobile platforms with an inference time of 1.36 ms and a memory footprint of 11.04 MB. CAPS ensures robust speech intelligibility in real-world settings, bridging the gap between efficiency and power savings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。