arXiv:2409.17069cs.SDcs.AI2024-09

用感知度量训练自编码器,能更好提取音乐特征用于流派分类。

The Effect of Perceptual Metrics on Music Representation Learning for Genre Classification

  • 用感知度量做损失函数训练自编码器,学习有意义的音乐表征。
  • 在音乐流派分类任务中,该方法比直接用度量作距离提升约5.2%准确率。
  • 适合对音乐表征学习和感知建模感兴趣的科研人员。

自然信号的主观质量可通过客观感知度量近似。设计用于模拟人类观察者感知行为的感知度量,通常反映了自然信号中的结构及神经通路特征。以感知度量作为损失函数训练的模型,可从这些度量所包含的结构中捕捉到具有感知意义的特征。本文证明,使用通过感知损失训练的自编码器提取的特征,在音乐理解任务(如流派分类)上的表现优于直接将这些度量用作分类器的距离。这一结果表明,采用感知度量作为表示学习的损失函数,可提升模型对新信号的泛化能力。

原文摘要 · Abstract (English)

The subjective quality of natural signals can be approximated with objective perceptual metrics. Designed to approximate the perceptual behaviour of human observers, perceptual metrics often reflect structures found in natural signals and neurological pathways. Models trained with perceptual metrics as loss functions can capture perceptually meaningful features from the structures held within these metrics. We demonstrate that using features extracted from autoencoders trained with perceptual losses can improve performance on music understanding tasks, i.e. genre classification, over using these metrics directly as distances when learning a classifier. This result suggests improved generalisation to novel signals when using perceptual metrics as loss functions for representation learning.

音乐表征感知度量自编码器流派分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。