自动化清理水下声学数据,让无标签音频也能训练出高效模型。
Automated data curation for self-supervised learning in underwater acoustic analysis
- 用船舶位置数据与水听器录音结合,自动筛选并平衡音频样本。
- 通过分层k均值聚类,从海量原始录音中提取多样且均衡的数据集。
- 适合做海洋生物监测和声污染评估的研究者使用。
海洋生态系统正面临声污染加剧的威胁,亟需监测以理解其变化与影响。被动声学监测(PAM)系统收集大量水下声音数据,但人工分析难以应对如此庞大的数据量,迫切需要自动化解决方案。尽管机器学习具备潜力,但多数水下声学数据缺乏标注。自监督学习在计算机视觉、自然语言处理和音频等领域已证明能从大规模无标签数据中学习,但其性能依赖于大规模、多样化且均衡的数据集。为此,本文提出一个完全自动化的自监督数据清洗流程,从美国水域的原始PAM数据中构建多样化且均衡的数据集。该流程整合了自动识别系统(AIS)数据与多个水听器的录音,利用分层k均值聚类对原始音频进行采样,并与AIS数据融合,生成平衡且多样的数据集。由此构建的高质量数据集可支持自监督学习模型开发,助力海洋哺乳动物监测与声污染评估等任务。
原文摘要 · Abstract (English)
The sustainability of the ocean ecosystem is threatened by increased levels of sound pollution, making monitoring crucial to understand its variability and impact. Passive acoustic monitoring (PAM) systems collect a large amount of underwater sound recordings, but the large volume of data makes manual analysis impossible, creating the need for automation. Although machine learning offers a potential solution, most underwater acoustic recordings are unlabeled. Self-supervised learning models have demonstrated success in learning from large-scale unlabeled data in various domains like computer vision, Natural Language Processing, and audio. However, these models require large, diverse, and balanced datasets for training in order to generalize well. To address this, a fully automated self-supervised data curation pipeline is proposed to create a diverse and balanced dataset from raw PAM data. It integrates Automatic Identification System (AIS) data with recordings from various hydrophones in the U.S. waters. Using hierarchical k-means clustering, the raw audio data is sampled and then combined with AIS samples to create a balanced and diverse dataset. The resulting curated dataset enables the development of self-supervised learning models, facilitating various tasks such as monitoring marine mammals and assessing sound pollution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。