修复音乐制作中失真的音源信号,还原原始音频质量。
Music Source Restoration
- 将混音建模为各音源独立失真后的叠加,而非简单相加。
- 构建了包含354小时未处理音源的RawStems数据集,支持8类主乐器分组。
- 提出首个适用于真实音乐生产流程的音源恢复基准方法U-Former。
我们提出音乐源恢复(MSR),填补理想化音源分离与真实音乐制作之间的空白。现有音乐源分离(MSS)方法假设混合信号仅为源信号的简单叠加,忽略了实际制作中常见的均衡、压缩、混响等信号退化。MSR将混合信号建模为各自退化后源信号的叠加,目标是恢复原始未退化信号。由于缺乏相关数据,我们构建了RawStems数据集,包含578首歌曲的未处理源信号,按8个主类和17个次级乐器组分类,总计354.13小时。据我们所知,这是首个包含层次化分类未处理音源的数据库。我们考虑频谱滤波、动态范围压缩、谐波失真、混响及有损编码作为可能的退化形式,并建立U-Former作为基线方法,验证了在该数据集上实现MSR的可行性。我们公开发布RawStems数据标注、退化模拟流程、训练代码及预训练模型。
原文摘要 · Abstract (English)
We introduce Music Source Restoration (MSR), a novel task addressing the gap between idealized source separation and real-world music production. Current Music Source Separation (MSS) approaches assume mixtures are simple sums of sources, ignoring signal degradations employed during music production like equalization, compression, and reverb. MSR models mixtures as degraded sums of individually degraded sources, with the goal of recovering original, undegraded signals. Due to the lack of data for MSR, we present RawStems, a dataset annotation of 578 songs with unprocessed source signals organized into 8 primary and 17 secondary instrument groups, totaling 354.13 hours. To the best of our knowledge, RawStems is the first dataset that contains unprocessed music stems with hierarchical categories. We consider spectral filtering, dynamic range compression, harmonic distortion, reverb and lossy codec as possible degradations, and establish U-Former as a baseline method, demonstrating the feasibility of MSR on our dataset. We release the RawStems dataset annotations, degradation simulation pipeline, training code and pre-trained models to be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。