arXiv:2507.09768cs.LGcs.SD2025-07中稿 · ICLR被引 2

让语音分离模型学会适时退出,动态节省计算资源。

Knowing When to Quit: Probabilistic Early Exits for Speech Separation

  • 设计可提前退出的神经网络架构,根据不确定性决定何时停止计算。
  • 在保持音质的前提下,测试时动态调节计算量,最高节省40%算力。
  • 适合资源受限设备如耳机、手机,且退出条件直观可解释。

近年来,基于深度学习的单通道语音分离性能显著提升,主要得益于更高效、参数更少的神经网络架构。然而,大多数架构固定了计算与参数预算,难以适应不同计算需求或资源环境,限制了其在移动设备和异构硬件(如耳机)中的应用。为此,本文设计了一种支持早期退出的语音分离与增强神经网络架构,并提出一种不确定性感知的概率框架,联合建模纯净语音信号与误差方差,据此推导出以期望信噪比为依据的随机早期退出条件。我们在语音分离与增强任务上进行了评估,结果表明引入早期退出不会损害重建质量;在变长音频上训练后,退出条件具有良好的校准性,在测试时动态调节计算量可实现显著的算力节约,且退出机制保持直接可解释。

原文摘要 · Abstract (English)

In recent years, deep learning-based single-channel speech separation has improved considerably, in large part driven by increasingly compute- and parameter-efficient neural network architectures. Most such architectures are, however, designed with a fixed compute and parameter budget and consequently cannot scale to varying compute demands or resources, which limits their use in embedded and heterogeneous devices such as mobile phones and hearables. To enable such use-cases we design a neural network architecture for speech separation and enhancement capable of early-exit, and we propose an uncertainty-aware probabilistic framework to jointly model the clean speech signal and error variance which we use to derive probabilistic early-exit conditions in terms of desired signal-to-noise ratios. We evaluate our methods on both speech separation and enhancement tasks where we demonstrate that early-exit capabilities can be introduced without compromising reconstruction, and that when trained on variable-length audio our early-exit conditions are well-calibrated and lead to considerable compute savings when used to dynamically scale compute at test time while remaining directly interpretable.

语音分离早期退出计算效率概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。