arXiv:2508.05878cs.SDcs.LG2025-08被引 1

用合成音频训练和弦识别模型,效果接近真实音乐数据。

Training chord recognition models on artificially generated audio

  • 用合成多轨音频(AAM)与真实数据混合训练Transformer模型
  • 在无真实数据时,仅用合成数据也能实现可接受的和弦识别准确率
  • 适合缺乏版权音乐数据的研究者使用

音乐信息检索中获取足够非版权音频用于模型训练和评估是一大难题。本研究对比了两种基于Transformer的神经网络模型,在音频中进行和弦序列识别,并检验了使用人工生成数据集的有效性。模型在多种组合下训练:人工多轨音频(AAM)、舒伯特《冬之旅》数据集、麦吉尔比尔榜单数据集,并采用根音、大小调和弦内容度量(CCM)三项指标进行评估。实验表明,尽管人工生成音乐与真人创作在复杂性和结构上存在差异,但前者在某些场景下仍具实用价值。具体而言,AAM可扩充较小的真实人类作曲数据集,或在无其他数据可用时,独立作为流行音乐和弦序列预测模型的训练集。

原文摘要 · Abstract (English)

One of the challenging problems in Music Information Retrieval is the acquisition of enough non-copyrighted audio recordings for model training and evaluation. This study compares two Transformer-based neural network models for chord sequence recognition in audio recordings and examines the effectiveness of using an artificially generated dataset for this purpose. The models are trained on various combinations of Artificial Audio Multitracks (AAM), Schubert's Winterreise Dataset, and the McGill Billboard Dataset and evaluated with three metrics: Root, MajMin and Chord Content Metric (CCM). The experiments prove that even though there are certainly differences in complexity and structure between artificially generated and human-composed music, the former can be useful in certain scenarios. Specifically, AAM can enrich a smaller training dataset of music composed by a human or can even be used as a standalone training set for a model that predicts chord sequences in pop music, if no other data is available.

和弦识别合成音频Transformer数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。