用符号序列实现弱监督打鼓转录,提升节奏准确性。
Weakly Supervised Tabla Stroke Transcription via TI-SDRM: A Rhythm-Aware Lattice Rescoring Framework
- 结合声学模型与节奏重打分框架,利用符号序列进行弱监督学习。
- 在真实和合成数据上显著降低打鼓错误率,提升转录精度。
- 适合研究印度古典音乐节奏分析与弱监督语音处理的学者。
Tabla 打击转录(TST)是分析印度古典音乐节奏结构的核心任务,但因复杂的节奏组织和强标注数据稀缺而极具挑战。现有方法多依赖于成本高昂且难以大规模获取的逐音符标注。本文在弱监督设置下解决 TST 问题,仅使用无时间对齐的符号打击序列。提出一种融合基于 CTC 的声学模型与序列级节奏重打分的框架。声学模型生成解码格网,再通过 extbf{$Tar{a}la$}-Independent Static--Dynamic Rhythmic Model (TI-SDRM) 进行优化,该模型通过自适应插值机制整合长期节奏结构与短期动态变化。构建了一个新的真实世界 tabla 独奏数据集及配套合成数据集,建立了首个面向印度古典音乐弱监督 TST 的基准。实验表明,相比仅使用声学模型的解码,本方法在打鼓错误率上实现持续且显著的降低,验证了显式节奏结构对准确转录的重要性。
原文摘要 · Abstract (English)
Tabla Stroke Transcription (TST) is central to the analysis of rhythmic structure in Hindustani classical music, yet remains challenging due to complex rhythmic organization and the scarcity of strongly annotated data. Existing approaches largely rely on fully supervised learning with onset-level annotations, which are costly and impractical at scale. This work addresses TST in a weakly supervised setting, using only symbolic stroke sequences without temporal alignment. We propose a framework that combines a CTC-based acoustic model with sequence-level rhythmic rescoring. The acoustic model produces a decoding lattice, which is refined using a \textbf{$T\bar{a}la$}-Independent Static--Dynamic Rhythmic Model (TI-SDRM) that integrates long-term rhythmic structure with short-term adaptive dynamics through an adaptive interpolation mechanism. We curate a new real-world tabla solo dataset and a complementary synthetic dataset, establishing the first benchmark for weakly supervised TST in Hindustani classical music. Experiments demonstrate consistent and substantial reductions in stroke error rate over acoustic-only decoding, confirming the importance of explicit rhythmic structure for accurate transcription.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。