研究BPE在单声与多声部音乐中的表现,发现其对乐器配置敏感且能捕捉抽象音乐特征。
Analyzing Byte-Pair Encoding on Monophonic and Polyphonic Symbolic Music: A Focus on Musical Phrase Segmentation
- 用BPE处理不同乐器配置的音乐数据,分析其子词构建机制
- 多声部音乐中BPE显著提升分句任务性能,单声部需特定合并次数才有效
- 适用于音乐信息检索、分句等任务,尤其适合复杂多声部场景
字节对编码(BPE)是自然语言处理中常用的子词词汇构建方法,近年来被应用于符号化音乐。由于符号音乐与文本存在显著差异,尤其是多声部音乐,本文研究了BPE在不同类型音乐内容下的行为表现。通过定性分析不同乐器配置下的BPE表现,并评估其在单声部与多声部音乐中的分句任务效果,结果表明:BPE训练过程高度依赖于乐器配置;其生成的“超标记”(supertokens)能有效捕捉抽象音乐内容。在分句任务中,BPE在多声部设置下显著提升性能,而在单声部音乐中仅在特定合并次数范围内产生增益。
原文摘要 · Abstract (English)
Byte-Pair Encoding (BPE) is an algorithm commonly used in Natural Language Processing to build a vocabulary of subwords, which has been recently applied to symbolic music. Given that symbolic music can differ significantly from text, particularly with polyphony, we investigate how BPE behaves with different types of musical content. This study provides a qualitative analysis of BPE's behavior across various instrumentations and evaluates its impact on a musical phrase segmentation task for both monophonic and polyphonic music. Our findings show that the BPE training process is highly dependent on the instrumentation and that BPE "supertokens" succeed in capturing abstract musical content. In a musical phrase segmentation task, BPE notably improves performance in a polyphonic setting, but enhances performance in monophonic tunes only within a specific range of BPE merges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。