提出新方法分离音乐音频的私有与共享特征,提升表示解耦性。
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
- 通过多视图自监督学习,显式分离私有与共享表征空间。
- 在受控实验中验证了特征解耦效果,优于现有方法。
- 适合需要可解释音乐表示的研究者和音频模型开发者。
自监督学习(SSL)为在无标注数据下学习鲁棒、通用的表示提供了强大途径。在音乐领域,由于标注数据稀缺,现有SSL方法通常依赖生成的监督信号和多视图冗余来构建预训练任务。然而,这些方法常导致表示纠缠并丢失视图特异性信息。本文提出一种新颖的自监督多视图音频学习框架,旨在激励私有与共享表示空间之间的分离。在受控环境下的音频解耦案例研究中,验证了该方法的有效性。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) offers a powerful way to learn robust, generalizable representations without labeled data. In music, where labeled data is scarce, existing SSL methods typically use generated supervision and multi-view redundancy to create pretext tasks. However, these approaches often produce entangled representations and lose view-specific information. We propose a novel self-supervised multi-view learning framework for audio designed to incentivize separation between private and shared representation spaces. A case study on audio disentanglement in a controlled setting demonstrates the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。