arXiv:2508.21153cs.SDcs.AI2025-08

轻量级音频修复模型,用压缩潜空间提升效率

WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration

  • 用神经音频编码器+潜空间扩散模型,降低计算开销
  • 在Voicebank+DEMAND上实现0.48~0.60的低谱距离
  • 适合资源受限场景,为后续优化打下基础

高质量音频在在线通信、虚拟助手和多媒体行业至关重要,但噪声、压缩和传输失真仍构成主要挑战。尽管扩散模型在音频修复中表现有效,但通常计算开销大,难以处理长段缺失。本文提出WaveLLDM(Wave Lightweight Latent Diffusion Model),将高效神经音频编码器与潜空间扩散模型结合,使音频在压缩潜空间中处理,降低复杂度同时保持重建质量。在Voicebank+DEMAND测试集上的实验表明,WaveLLDM实现了0.48至0.60的低对数谱距离(LSD)得分,并具备良好泛化能力。但在感知质量和语音清晰度方面仍逊于先进方法,整体客观评分WB-PESQ为1.62~1.71,STOI为0.76~0.78,归因于架构调优不足、缺乏微调及训练时间有限。其灵活架构为未来研究提供了坚实基础。

原文摘要 · Abstract (English)

High-quality audio is essential in a wide range of applications, including online communication, virtual assistants, and the multimedia industry. However, degradation caused by noise, compression, and transmission artifacts remains a major challenge. While diffusion models have proven effective for audio restoration, they typically require significant computational resources and struggle to handle longer missing segments. This study introduces WaveLLDM (Wave Lightweight Latent Diffusion Model), an architecture that integrates an efficient neural audio codec with latent diffusion for audio restoration and denoising. Unlike conventional approaches that operate in the time or spectral domain, WaveLLDM processes audio in a compressed latent space, reducing computational complexity while preserving reconstruction quality. Empirical evaluations on the Voicebank+DEMAND test set demonstrate that WaveLLDM achieves accurate spectral reconstruction with low Log-Spectral Distance (LSD) scores (0.48 to 0.60) and good adaptability to unseen data. However, it still underperforms compared to state-of-the-art methods in terms of perceptual quality and speech clarity, with WB-PESQ scores ranging from 1.62 to 1.71 and STOI scores between 0.76 and 0.78. These limitations are attributed to suboptimal architectural tuning, the absence of fine-tuning, and insufficient training duration. Nevertheless, the flexible architecture that combines a neural audio codec and latent diffusion model provides a strong foundation for future development.

音频修复扩散模型轻量级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。