arXiv:2608.07423cs.SDcs.LG2026-08中稿 · Interspeech 2026

云端协作提升低算力设备语音增强效果

Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement

论文配图:Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement
图 1 · 摘自论文原文
  • 用服务器延迟输出和中间特征指导边缘端推理
  • 多通道维纳滤波融合云端与本地协方差矩阵
  • 在极低额外开销下显著超越纯本地模型

低延迟、低算力的语音增强对具有实时通信需求的可穿戴设备至关重要,但严格的计算约束显著限制了本地性能。知识增强已被提出作为通过更强大的服务器端模型提升边缘模型性能的有效方法,但在语音增强任务中性能提升有限。本文提出一种协作框架,包含三项技术:(1) 将服务器延迟输出作为额外输入;(2) 层级特征增强,将服务器的中间表示传递给边缘端以指导推理;(3) 协作式多通道维纳滤波,融合由服务器和边缘模型分别估计的加权协方差矩阵以改善波束成形。实验结果表明,所提出的协作框架在仅增加极少计算开销的前提下,显著优于纯边缘端基线模型。

原文摘要 · Abstract (English)

Low-latency, low-compute speech enhancement is essential for wearable devices with real-time communication requirements, but strict computational constraints significantly limit on-device performance. Knowledge Boosting has been proposed as an effective approach to improve edge model performance by leveraging a more capable server-side model, but performance gains for speech enhancement have been limited. We propose a collaborative framework incorporating three techniques: (1) delayed server output as additional input, (2) layerwise feature boosting that transfers intermediate server representations to guide edge inference, and (3) collaborative multichannel Wiener filtering, which fuses weighted covariance matrices estimated from both server and edge models for improved beamforming. Experimental results demonstrate that the proposed collaborative framework significantly outperforms the edge-only baseline with minimal additional computational overhead.

语音增强边缘计算协同学习低算力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。