用符号节奏规律提升塔布拉鼓转录准确率
Weakly Supervised Tabla Stroke Transcription via an Adaptive Dynamic Rhythm Language Model (ADRM)

- 结合声学模型与自适应节奏语言模型进行解码重评分
- 在无时间对齐标注下,击打错误率显著降低
- 适合研究印度音乐节奏分析与弱监督学习者
塔布拉鼓击打转录(TST)是分析北印度音乐节奏结构的核心任务,但因节奏组织复杂多变且强标注数据稀缺而极具挑战。现有方法主要依赖逐音符标注的全监督学习,成本高昂且难以规模化。本文提出一种弱监督框架,仅使用符号化击打序列,无需起始时间对齐。通过连接时序分类(CTC)声学模型生成解码路径,再利用自适应动态节奏语言模型(ADRM)进行重评分。ADRM融合了节拍(tāla)条件下的符号化节奏规律与局部击打动态特性。此外,本文发布新数据集《塔布拉即兴演奏数据集》及配套合成数据集,支持序列级弱监督转录。实验表明,相较于仅使用声学模型的解码,ADRM在多个指标上均显著降低击打错误率,验证了在解码重评分中引入符号化节奏规律的有效性。
原文摘要 · Abstract (English)
Tabla Stroke Transcription (TST) is central to the analysis of rhythmic structure in Hindustani music, yet it remains challenging due to complex and dynamic rhythmic organization and the scarcity of strongly annotated data. Existing approaches largely rely on fully supervised learning with onset-level annotations, which are costly and impractical at scale. This work addresses TST in a weakly supervised setting, using only symbolic stroke sequences without temporal alignment of onsets. We propose a framework that combines a Connectionist Temporal Classification (CTC)-based acoustic model with a sequence-level rhythmic language model for rescoring, similar to that used in automatic speech recognition. The acoustic model produces a decoding lattice, which is refined using an Adaptive Dynamic Rhythm Language Model (ADRM) that combines $t\bar{a}la$-conditioned symbolic rhythmic regularities with local stroke dynamics. Moreover, we release a new performance-recorded tabla dataset, named \emph{Tabla Improvisation Dataset}, along with a complementary synthetic dataset for sequence-level weakly supervised TST. Experiments demonstrate consistent and substantial reductions in stroke error rates with ADRM compared to those with acoustic-only decoding, confirming the benefit of incorporating symbolic rhythmic regularities during lattice rescoring for accurate transcription.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。