用RQ-VAE对乐段分块编码,让机器更懂音乐理论和情绪。
MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
- 基于Transformer的分小节量化编码,生成高保真音乐代码。
- 在和弦识别与情绪判断上超越现有方法,生成质量相当。
- 能捕捉音乐理论概念,适合音乐生成与理解研究者使用。
离散表示学习在图像、语音和语言等领域已取得显著进展。受此启发,我们提出MuseTok,一种面向符号音乐的分块量化方法,并在音乐生成与语义理解任务中验证其有效性。MuseTok在基于Transformer的编码器-解码器框架下,对小节级音乐片段采用残差向量量化变分自编码器(RQ-VAE)进行建模,生成的音乐代码可实现高保真重建,并准确理解音乐理论。为全面评估,我们在旋律提取、和弦识别及情绪识别等任务上测试MuseTok。引入该表示的模型在语义理解上优于以往基线,同时在内容生成性能上保持相当。此外,通过真实标签与合成数据的定性分析显示,MuseTok能从大规模音乐数据中有效捕获底层音乐概念。
原文摘要 · Abstract (English)
Discrete representation learning has shown promising results across various domains, including generation and understanding in image, speech and language. Inspired by these advances, we propose MuseTok, a tokenization method for symbolic music, and investigate its effectiveness in both music generation and understanding tasks. MuseTok employs the residual vector quantized-variational autoencoder (RQ-VAE) on bar-wise music segments within a Transformer-based encoder-decoder framework, producing music codes that achieve high-fidelity music reconstruction and accurate understanding of music theory. For comprehensive evaluation, we apply MuseTok to music generation and semantic understanding tasks, including melody extraction, chord recognition, and emotion recognition. Models incorporating MuseTok outperform previous representation learning baselines in semantic understanding while maintaining comparable performance in content generation. Furthermore, qualitative analyses on MuseTok codes, using ground-truth categories and synthetic datasets, reveal that MuseTok effectively captures underlying musical concepts from large music collections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。