用轻量注意力U-Net自动清除录音中的呼吸声,提升音质效率。
Attention-Based Efficient Breath Sound Removal in Studio Audio Recordings
- 基于注意力U-Net架构,仅需190万参数
- 训练仅3.2小时,精度显著优于现有模型
- 适合音频工程师快速处理专业录音
本研究提出一种参数高效模型,利用注意力U-Net架构自动检测并消除人声录音中的非语音声音,特别是呼吸声。该任务在音频工程领域至关重要,但长期未受充分关注。传统人工处理依赖经验且耗时极长,现有自动化方法在效率与精度上均表现不足。所提模型通过深度学习技术实现流程简化与精度提升,采用源自DAPS的数据集,以对数频谱图为输入,并引入早停机制防止过拟合。模型仅需190万参数与3.2小时训练时间,即可达到与顶尖模型相当的输出效果,显著提升音频制作质量与一致性,为行业带来重要突破。
原文摘要 · Abstract (English)
In this research, we present an innovative, parameter-efficient model that utilizes the attention U-Net architecture for the automatic detection and eradication of non-speech vocal sounds, specifically breath sounds, in vocal recordings. This task is of paramount importance in the field of sound engineering, despite being relatively under-explored. The conventional manual process for detecting and eliminating these sounds requires significant expertise and is extremely time-intensive. Existing automated detection and removal methods often fall short in terms of efficiency and precision. Our proposed model addresses these limitations by offering a streamlined process and superior accuracy, achieved through the application of advanced deep learning techniques. A unique dataset, derived from Device and Produced Speech (DAPS), was employed for this purpose. The training phase of the model emphasizes a log spectrogram and integrates an early stopping mechanism to prevent overfitting. Our model not only conserves precious time for sound engineers but also enhances the quality and consistency of audio production. This constitutes a significant breakthrough, as evidenced by its comparative efficiency, necessitating only 1.9M parameters and a training duration of 3.2 hours - markedly less than the top-performing models in this domain. The model is capable of generating identical outputs as previous models with drastically improved precision, making it an optimal choice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。