arXiv:2411.02711cs.SDeess.AS2024-11被引 3

提出新方法分离音乐音频的私有与共享特征,提升表示解耦性。

Self-Supervised Multi-View Learning for Disentangled Music Audio Representations

  • 通过多视图自监督学习,显式分离私有与共享表征空间。
  • 在受控实验中验证了特征解耦效果,优于现有方法。
  • 适合需要可解释音乐表示的研究者和音频模型开发者。

自监督学习(SSL)为在无标注数据下学习鲁棒、通用的表示提供了强大途径。在音乐领域,由于标注数据稀缺,现有SSL方法通常依赖生成的监督信号和多视图冗余来构建预训练任务。然而,这些方法常导致表示纠缠并丢失视图特异性信息。本文提出一种新颖的自监督多视图音频学习框架,旨在激励私有与共享表示空间之间的分离。在受控环境下的音频解耦案例研究中,验证了该方法的有效性。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) offers a powerful way to learn robust, generalizable representations without labeled data. In music, where labeled data is scarce, existing SSL methods typically use generated supervision and multi-view redundancy to create pretext tasks. However, these approaches often produce entangled representations and lose view-specific information. We propose a novel self-supervised multi-view learning framework for audio designed to incentivize separation between private and shared representation spaces. A case study on audio disentanglement in a controlled setting demonstrates the effectiveness of our method.

自监督学习音频表示特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。