构建首个印度古典唱腔装饰音数据集,提升声乐装饰音识别精度。
Recognizing Ornaments in Vocal Indian Art Music with Active Annotation
- 用专家标注的互动工具标记六类唱腔装饰音事件。
- 在长音频切片中保持装饰音边界,提升检测精度。
- 在自建与外部数据集上均优于基线模型,适合音乐信息检索研究。
装饰音、润腔或微调音是多种音乐传统中旋律表达的核心元素,为表演增添深度、细腻感与情感冲击。识别歌唱中的装饰音对音乐信息检索(MIR)至关重要,具有音乐教学、歌手识别、流派分类及可控歌声生成等应用前景。然而,缺乏标注数据集和专用建模方法仍是该领域发展的主要障碍。本文提出Rāga Ornamentation Detection(ROD)数据集,由专家音乐家精选并标注的印度古典音乐录音组成,采用定制的人机协作工具对六类声乐装饰音进行事件级标注。基于此数据集,我们构建了基于深度时序分析的装饰音检测模型,能够在长音频切片中保留装饰音边界。通过在ROD数据集内采用不同训练-测试配置,并在另一组人工标注的印度古典音乐会录音上评估,实验结果表明所提方法显著优于基线CRNN模型。
原文摘要 · Abstract (English)
Ornamentations, embellishments, or microtonal inflections are essential to melodic expression across many musical traditions, adding depth, nuance, and emotional impact to performances. Recognizing ornamentations in singing voices is key to MIR, with potential applications in music pedagogy, singer identification, genre classification, and controlled singing voice generation. However, the lack of annotated datasets and specialized modeling approaches remains a major obstacle for progress in this research area. In this work, we introduce Rāga Ornamentation Detection (ROD), a novel dataset comprising Indian classical music recordings curated by expert musicians. The dataset is annotated using a custom Human-in-the-Loop tool for six vocal ornaments marked as event-based labels. Using this dataset, we develop an ornamentation detection model based on deep time-series analysis, preserving ornament boundaries during the chunking of long audio recordings. We conduct experiments using different train-test configurations within the ROD dataset and also evaluate our approach on a separate, manually annotated dataset of Indian classical concert recordings. Our experimental results support the superior performance of our proposed approach over the baseline CRNN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。