用音程替代绝对音高,让音乐模型更懂旋律走势
Evaluating Interval-based Tokenization for Pitch Representation in Symbolic Music Analysis
- 用音程构建符号音乐的分词方式,替代传统绝对音高
- 在三个音乐分析任务中,性能均优于传统方法
- 提升模型可解释性,适合关注旋律结构的研究者
符号音乐分析常采用为自然语言处理设计的模型(如Transformer),需将数据转化为序列,依赖分词。现有音乐分词多使用绝对MIDI值表示音高,但音乐研究强调旋律轮廓和和声关系等高层表示,音程比绝对音高更具表现力。本文提出一种通用的基于音程的分词框架,在三个音乐分析任务上验证其有效性:不仅提升了模型性能,还增强了可解释性。
原文摘要 · Abstract (English)
Symbolic music analysis tasks are often performed by models originally developed for Natural Language Processing, such as Transformers. Such models require the input data to be represented as sequences, which is achieved through a process of tokenization. Tokenization strategies for symbolic music often rely on absolute MIDI values to represent pitch information. However, music research largely promotes the benefit of higher-level representations such as melodic contour and harmonic relations for which pitch intervals turn out to be more expressive than absolute pitches. In this work, we introduce a general framework for building interval-based tokenizations. By evaluating these tokenizations on three music analysis tasks, we show that such interval-based tokenizations improve model performances and facilitate their explainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。