arXiv:2608.10979cs.SD2026-08中稿 · ISMIR 2026

用VQ-VAE学习韩国民乐的连续音高模式,无需标注即可捕捉核心音乐特征。

Pitch Contour Tokenization using VQ-VAE and Its Application on Korean Traditional Music Analysis

论文配图:Pitch Contour Tokenization using VQ-VAE and Its Application on Korean Traditional Music Analysis
图 1 · 摘自论文原文
  • 通过VQ-VAE从无标注音频中自动学习局部音高轮廓的离散表示。
  • 在韩国传统音乐上,模型学习到的符号可无监督还原专家定义的四声类别。
  • 适用于以音高轮廓为核心的音乐分析,适合民族音乐学与跨文化研究者。

音乐的计算分析常依赖离散表示,但许多音乐传统以连续音高运动为核心,难以分割为音符类单元。本文提出一种方法,使用VQ-VAE从无标注音频中直接学习局部音高轮廓模式的词汇表,将固定长度的轮廓片段量化为有限码本。为使学习到的符号对分段位置、时间微调和音域变化具有稳定性,训练时采用在一组候选时移与音域变换中最佳对齐下的重构目标。该方法应用于韩国民间唱曲(Pansori),学习到的符号能无监督地恢复专家定义的四声类别(sigimsae),且个体符号与主要两种调式(Gyemyeonjo与Ujo)高度对应,支持其作为轮廓主导传统音乐的语料级分析单位。

原文摘要 · Abstract (English)

Computational analysis of music often relies on discrete representations, yet many musical traditions are organized around continuous pitch movement that resists segmentation into note-like units. For such traditions, the discrete units that analysis would build on are not given in advance. We address this gap by learning a vocabulary of local pitch-contour patterns directly from unlabeled audio, using a VQ-VAE that quantizes fixed-length contour segments into a finite codebook. To make the learned tokens stable across segmentation positions and small variations in timing and pitch range, we train the model with a reconstruction objective evaluated under the best alignment among a set of candidate temporal and pitch-domain transformations. Applied to Korean traditional music, the learned tokens recover information about expert-defined sigimsae categories without supervision, and in pansori individual tokens align with the two principal modes, Gyemyeonjo and Ujo, supporting their use as units for corpus-level analysis of contour-centric traditions.

音高轮廓VQ-VAE韩国民乐无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。