arXiv:2608.23038cs.SD2026-08

提出可保证收敛的音频信号恢复神经网络,解决传统方法无法控制稳定性的难题。

LipsAM: Lipschitz-continuous Neural Networks for Convergent Plug-and-Play Audio Signal Recovery

  • 构建仅处理复数信号幅度的神经网络,理论证明其 Lipschitz 连续性
  • 设计 LipsAM 架构并精确计算其 Lipschitz 常数,适用于时频掩码等常见结构
  • 集成到插件式信号恢复框架中,实现语音去混响的稳定收敛

深度神经网络(DNN)的 Lipschitz 连续性对于建立其行为的理论保障至关重要。尽管已有多种方法用于构建 Lipschitz 连续架构并控制其常数,但音频信号处理中常见的若干 DNN 架构仍超出现有理论框架,阻碍了声学应用中的 Lipschitz 模型发展。特别是,尽管广泛使用,仅分别处理复数信号幅度和相位的 DNN 在现有框架下无法实现 Lipschitz 连续。为此,本文建立了幅度调节器(AMs)的理论基础,这类 DNN 仅作用于复数输入的幅度,且可证明 Lipschitz 连续。我们推导出 AM 的必要充分条件,并提出了对应于常见音频架构(如时频掩码)的 LipsAMs。此外,开发了高效评估 Lipschitz 常数的框架,并对部分架构进行了解析推导。作为应用,提出 CoReM-LipsAM(基于 LipsAM 的可控残差映射),将 DNN 作为数据驱动先验嵌入模型基信号处理算法中。该 PnP 算法的收敛性由 CoReM-LipsAM 架构保证,并通过语音去混响实验得到实证验证。

原文摘要 · Abstract (English)

The Lipschitz continuity of deep neural networks (DNNs) is essential for establishing theoretical guarantees regarding their behavior. From both theoretical and practical perspectives, various methods have been proposed to construct Lipschitz-continuous architectures and control their Lipschitz constants. However, several DNN architectures common in audio signal processing fall outside the scope of existing theoretical frameworks, hindering the development of Lipschitz-continuous models in acoustic applications. In particular, despite their widespread adoption, DNNs that separately process the magnitude and phase of complex-valued signals cannot be Lipschitz continuous under existing frameworks. In this paper, to address this limitation, we establish a theoretical foundation for constructing amplitude modifiers (AMs), a class of DNN architectures that operate solely on the magnitude of a complex-valued input, with provable Lipschitz continuity. Specifically, we derive a necessary and sufficient condition for an AM to be Lipschitz continuous and propose LipsAMs (Lipschitz-continuous AMs) corresponding to common architectures for audio signals, including time-frequency masking. Furthermore, we develop an efficient framework for evaluating their Lipschitz constants and analytically derive these constants for some of the proposed architectures. As an application, we propose CoReM-LipsAM (Controlled Residual Maps via LipsAM) for plug-and-play (PnP) audio signal recovery, integrating a DNN as a data-driven prior within a model-based signal processing algorithm. The convergence of the obtained PnP algorithm is structurally guaranteed by the CoReM-LipsAM architecture and empirically validated through speech dereverberation experiments.

音频恢复Lipschitz神经网络收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。