arXiv:2603.21684cs.SDcs.LG2026-03中稿 · IEEE ICASSP 2026被引 1

提出 LipsAM,让语音处理模型更稳定可靠。

LipsAM: Lipschitz-Continuous Amplitude Modifier for Audio Signal Processing and its Application to Plug-and-Play Dereverberation

  • 设计满足李普希茨连续性的语音幅度调节模块。
  • 在去混响任务中验证了模型稳定性提升。
  • 适合关注模型可靠性的语音处理研究者。

深度神经网络(DNN)的鲁棒性可通过其李普希茨连续性进行认证,构建李普希茨连续的DNN是当前活跃的研究方向。然而,由于与现有成果兼容性差,音频处理领域的DNN尚未成为重点。本文聚焦于处理音频信号的常用架构——幅度调节器(AM),提出了其李普希茨连续变体,称为LipsAM。我们证明了AM具备李普希茨连续性的充分条件,并给出了两种LipsAM架构作为实例。将所提架构应用于即插即用的语音去混响算法,在数值实验中验证了其增强的稳定性。

原文摘要 · Abstract (English)

The robustness of deep neural networks (DNNs) can be certified through their Lipschitz continuity, which has made the construction of Lipschitz-continuous DNNs an active research field. However, DNNs for audio processing have not been a major focus due to their poor compatibility with existing results. In this paper, we consider the amplitude modifier (AM), a popular architecture for handling audio signals, and propose its Lipschitz-continuous variants, which we refer to as LipsAM. We prove a sufficient condition for an AM to be Lipschitz continuous and propose two architectures as examples of LipsAM. The proposed architectures were applied to a Plug-and-Play algorithm for speech dereverberation, and their improved stability is demonstrated through numerical experiments.

语音处理李普希茨去混响神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。