arXiv:2608.06633cs.SD2026-08中稿 · the 27th Internati…被引 1

用多模态音频表示识别韩国传统唱腔的调式,避免模型偷懒记忆曲目。

Frame-Level Pansori Mode Classification with Complementary Audio Representations

论文配图:Frame-Level Pansori Mode Classification with Complementary Audio Representations
图 1 · 摘自论文原文
  • 融合声谱、音高轮廓、钢琴卷帘和跨文化预训练特征进行调式分类
  • 在完整作品保留测试中,三类主要调式的F1值仅下降2.1-3.6分
  • 揭示了打击乐对changjo调式的关键影响及通用预训练在Ujo/Gyemyeonjo上的失败

Pansori是韩国民谣的一种传统声乐形式,其调式(jo)不仅由音阶决定,更依赖于音高集合、微音装饰(sigimsae)与嗓音音色的复杂交织。本研究构建了46小时的逐帧调式标注数据集,由专家标注所有五种经典板腔(batang),并评估四种互补输入表示(梅尔频谱图、基频轮廓、MIDI钢琴卷帘、多文化预训练自监督编码器)在两种防捷径学习的分割策略下的表现。在三个代表性调式上,当完整作品被排除时,模型性能仅下降2.1–3.6个F1点,表明模型学习的是调式相关特征而非记忆曲目。每类结果进一步显示,源分离会移除changjo依赖的打击乐线索,而通用多文化预训练在Ujo与Gyemyeonjo区分上表现不佳。跨模态不一致的定性分析恢复了音乐学已知现象,并与现代changjak pansori的乐谱分析结论一致。

原文摘要 · Abstract (English)

Pansori is a traditional Korean vocal genre whose mode system (jo) is defined not by scale alone but by the entanglement of pitch collection, microtonal ornament (sigimsae), and vocal timbre. In this study, we introduce a 46-hour frame-level pansori mode annotation, expert-labeled across all five canonical batang, and evaluate four complementary input representations (mel spectrogram, F0 contour, MIDI piano roll, and a multi-cultural SSL encoder) under two split strategies designed to detect shortcut learning. Across the three well-represented modes, performance degrades by only 2.1--3.6 points of F1 when entire works are held out, indicating that the models learn mode-relevant features rather than memorizing repertoire. Per-class results further show that source separation removes the percussion cue on which changjo depends, and that generic multi-cultural pre-training fails specifically on the Ujo--Gyemyeonjo distinction. Qualitative analysis of cross-modal disagreement recovers musicologically documented phenomena and agrees with published score-based analyses of modern changjak pansori.

音乐信息检索调式识别多模态学习传统音乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。