融合扩散模型与GAN,提升异常声音检测精度和定位能力。
TLDiffGAN: A Latent Diffusion-GAN Framework with Temporal Information Fusion for Anomalous Sound Detection
- 用潜空间扩散模型增强GAN生成器,提升样本质量。
- 在DCASE 2020数据集上达到最优检测性能,定位更精准。
- 适合关注音频异常检测与时间模式捕捉的研究者。
现有无监督异常声音检测的生成模型难以充分捕捉正常声音的复杂特征分布,而强大的扩散模型在此领域的潜力尚未被充分挖掘。为此,我们提出TLDiffGAN框架,包含两个互补分支:一是在GAN生成器中引入潜空间扩散模型进行对抗训练,使判别器面临更大挑战,提升生成样本质量;二是利用预训练音频模型编码器直接从原始波形提取特征,用于辅助判别。该框架有效捕获了来自原始音频和梅尔频谱图的正常声音特征表示。此外,引入一种TMixup频谱增强技术,提升对细微、局部时间模式的敏感性。在DCASE 2020 Challenge Task 2数据集上的大量实验表明,TLDiffGAN表现出卓越的检测性能,并具备强异常时频定位能力。
原文摘要 · Abstract (English)
Existing generative models for unsupervised anomalous sound detection are limited by their inability to fully capture the complex feature distribution of normal sounds, while the potential of powerful diffusion models in this domain remains largely unexplored. To address this challenge, we propose a novel framework, TLDiffGAN, which consists of two complementary branches. One branch incorporates a latent diffusion model into the GAN generator for adversarial training, thereby making the discriminator's task more challenging and improving the quality of generated samples. The other branch leverages pretrained audio model encoders to extract features directly from raw audio waveforms for auxiliary discrimination. This framework effectively captures feature representations of normal sounds from both raw audio and Mel spectrograms. Moreover, we introduce a TMixup spectrogram augmentation technique to enhance sensitivity to subtle and localized temporal patterns that are often overlooked. Extensive experiments on the DCASE 2020 Challenge Task 2 dataset demonstrate the superior detection performance of TLDiffGAN, as well as its strong capability in anomalous time-frequency localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。