让模型学会区分音乐中每种乐器的相似性,更贴近人耳感受。
Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference
- 分步处理音源分离与特征提取,端到端微调减少误差影响。
- 多任务学习让模型更好分离不同乐器的特征表示。
- 结合人类听感偏好训练,提升结果与人感知的一致性。
本文提出基于单个乐器声音的音乐相似性表征学习(InMSRL),利用音乐源分离(MSS)和人类偏好,在推理阶段无需纯净乐器音轨。提出三种方法:首先,针对级联方法引入端到端微调(E2E-FT),使模型在音源分离后仍能有效提取特征,减轻分离误差的影响;其次,针对直接方法提出多任务学习,通过重建与解耦特征的联合优化,增强乐器特征的解耦能力;第三,采用感知对齐微调(PAFT),利用人类听觉偏好指导训练,使模型学习更符合人类感知的音乐相似性。实验表明:1)级联方法使用E2E-FT显著提升性能;2)直接方法的多任务学习有助于提高特征解耦效果;3)PAFT显著增强感知一致性表现;4)结合E2E-FT与PAFT的级联方案优于使用多任务学习与PAFT的直接方案。
原文摘要 · Abstract (English)
This paper proposes music similarity representation learning (MSRL) based on individual instrument sounds (InMSRL) utilizing music source separation (MSS) and human preference without requiring clean instrument sounds during inference. We propose three methods that effectively improve performance. First, we introduce end-to-end fine-tuning (E2E-FT) for the Cascade approach that sequentially performs MSS and music similarity feature extraction. E2E-FT allows the model to minimize the adverse effects of a separation error on the feature extraction. Second, we propose multi-task learning for the Direct approach that directly extracts disentangled music similarity features using a single music similarity feature extractor. Multi-task learning, which is based on the disentangled music similarity feature extraction and MSS based on reconstruction with disentangled music similarity features, further enhances instrument feature disentanglement. Third, we employ perception-aware fine-tuning (PAFT). PAFT utilizes human preference, allowing the model to perform InMSRL aligned with human perceptual similarity. We conduct experimental evaluations and demonstrate that 1) E2E-FT for Cascade significantly improves InMSRL performance, 2) the multi-task learning for Direct is also helpful to improve disentanglement performance in the feature extraction, 3) PAFT significantly enhances the perceptual InMSRL performance, and 4) Cascade with E2E-FT and PAFT outperforms Direct with the multi-task learning and PAFT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。