用非线性动力系统建模语音混沌特性,大幅压缩判别器参数量。
NLDSI-BWE: Non Linear Dynamical Systems-Inspired Multi Resolution Discriminators for Speech Bandwidth Extension
- 基于递归与李雅普诺夫指数设计双判别器,捕捉语音内在混沌动态。
- 判别器参数量减少44倍(约2200万→480万),性能超越现有模型。
- 适合关注语音生成、小规模高效判别器的学者和工程师。
本文设计了两种受非线性动力系统启发的判别器——多尺度递归判别器(MSRD)和多分辨率李雅普诺夫判别器(MRLD),以显式建模语音固有的确定性混沌特性。MSRD基于递归图表示,捕捉语音的自相似动态;MRLD基于李雅普诺夫指数,刻画非线性波动及对初始条件的敏感性。通过深度可分离卷积优化结构设计,所提框架在判别器参数量上相比先前的AP-BWE模型实现44倍压缩(约2200万 → 约480万)。据我们所知,这是首次将语音声带产生中的细微非线性混沌物理特性用于监督带宽扩展任务,显著降低判别器规模。
原文摘要 · Abstract (English)
In this paper, we design two nonlinear dynamical systems-inspired discriminators -- the Multi-Scale Recurrence Discriminator (MSRD) and the Multi-Resolution Lyapunov Discriminator (MRLD) -- to \textit{explicitly} model the inherent deterministic chaos of speech. MSRD is designed based on Recurrence representations to capture self-similarity dynamics. MRLD is designed based on Lyapunov exponents to capture nonlinear fluctuations and sensitivity to initial conditions. Through extensive design optimization and the use of depthwise-separable convolutions in the discriminators, our framework surpasses prior AP-BWE model with a 44x reduction in the discriminator parameter count \textbf{($\sim$ 22M vs $\sim$ 0.48M)}. To the best of our knowledge, for the first time, this paper demonstrates how BWE can be supervised by the subtle non-linear chaotic physics of voiced sound production to achieve a significant reduction in the discriminator size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。