arXiv:2410.21557cs.SDcs.LG2024-10

无真值数据下用新方法从噪声谱图中提取目标声学特征。

A Novel Score-CAM based Denoiser for Spectrographic Signature Extraction without Ground Truth

  • 基于Score-CAM设计生成对抗网络,无真值训练生成清晰谱图。
  • 在无标签数据上实现顶尖去噪效果,分类准确率超越现有标准。
  • 适用于真实世界多类声学数据,具强泛化能力,适合工业级应用。

基于声纳的音频分类技术是水下声学领域的研究热点。被动声纳传感器捕获的水下噪声包含各类传播信号,转化为谱图后常混杂大量干扰频段,导致目标信号难以辨识。多数从水下音频中提取的谱图因杂波过多而无法使用,且缺乏清洁的标注数据,严重制约了分类模型的训练。本文提出一种新型无真值的Score-CAM去噪器,可从噪声谱图中提取目标声学特征。通过构建新型生成对抗网络架构,学习并生成与低特征输入分布相似的合成谱图数据;同时提出通用化的类激活映射去噪模块,适用于多种声学数据分布,包括真实世界数据。实验表明,该方法在去噪精度和分类性能上均达到当前最优水平,不仅适用于音频数据,还可推广至全球范围内各类机器学习数据分布。

原文摘要 · Abstract (English)

Sonar based audio classification techniques are a growing area of research in the field of underwater acoustics. Usually, underwater noise picked up by passive sonar transducers contains all types of signals that travel through the ocean and is transformed into spectrographic images. As a result, the corresponding spectrograms intended to display the temporal-frequency data of a certain object often include the tonal regions of abundant extraneous noise that can effectively interfere with a 'contact'. So, a majority of spectrographic samples extracted from underwater audio signals are rendered unusable due to their clutter and lack the required indistinguishability between different objects. With limited clean true data for supervised training, creating classification models for these audio signals is severely bottlenecked. This paper derives several new techniques to combat this problem by developing a novel Score-CAM based denoiser to extract an object's signature from noisy spectrographic data without being given any ground truth data. In particular, this paper proposes a novel generative adversarial network architecture for learning and producing spectrographic training data in similar distributions to low-feature spectrogram inputs. In addition, this paper also a generalizable class activation mapping based denoiser for different distributions of acoustic data, even real-world data distributions. Utilizing these novel architectures and proposed denoising techniques, these experiments demonstrate state-of-the-art noise reduction accuracy and improved classification accuracy than current audio classification standards. As such, this approach has applications not only to audio data but for countless data distributions used all around the world for machine learning.

声学识别无监督学习谱图去噪GAN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。