让耳机在低功耗下实时处理语音,提升降噪与音质。
Wireless Hearables With Programmable Speech AI Accelerators
- 自研硬件加速器+优化神经网络,实现低延迟语音增强。
- 每6毫秒音频块处理仅需5.54毫秒,功耗71.6毫瓦。
- 适合对本地语音处理有高要求的智能耳机开发者。
传统观点认为,在功耗受限的微型无线助听设备上部署持续运行的深度学习语音模型极具挑战性,因其对计算和输入输出有严苛要求。本文提出NeuralAids,一种全本地化语音人工智能系统,可在紧凑、电池供电的设备上实现实时语音增强与降噪。通过三项关键技术突破:1)集成语音AI加速器的无线助听平台,实现高效本地流式推理;2)专为低延迟、高质量设计的双路径神经网络;3)软硬件协同设计,采用混合精度量化与量化感知训练,在严格功耗约束下达成实时性能。系统以6毫秒音频块为单位实时处理,推理时间仅5.54毫秒,功耗71.6毫瓦。真实场景评估(含28名用户)显示,其在语音质量与噪声抑制方面优于现有本地模型,为下一代全本地化智能助听设备铺平道路。
原文摘要 · Abstract (English)
The conventional wisdom has been that designing ultra-compact, battery-constrained wireless hearables with on-device speech AI models is challenging due to the high computational demands of streaming deep learning models. Speech AI models require continuous, real-time audio processing, imposing strict computational and I/O constraints. We present NeuralAids, a fully on-device speech AI system for wireless hearables, enabling real-time speech enhancement and denoising on compact, battery-constrained devices. Our system bridges the gap between state-of-the-art deep learning for speech enhancement and low-power AI hardware by making three key technical contributions: 1) a wireless hearable platform integrating a speech AI accelerator for efficient on-device streaming inference, 2) an optimized dual-path neural network designed for low-latency, high-quality speech enhancement, and 3) a hardware-software co-design that uses mixed-precision quantization and quantization-aware training to achieve real-time performance under strict power constraints. Our system processes 6 ms audio chunks in real-time, achieving an inference time of 5.54 ms while consuming 71.6 mW. In real-world evaluations, including a user study with 28 participants, our system outperforms prior on-device models in speech quality and noise suppression, paving the way for next-generation intelligent wireless hearables that can enhance hearing entirely on-device.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。