用少量标签自动标注印度音乐数据,提升标注效率与质量。
Learning from Limited Labels: Transductive Graph Label Propagation for Indian Music Analysis
- 构建音频嵌入相似图,通过图传播将少量标签扩展到大量无标签数据。
- 在拉格识别和乐器分类任务中,标签减少90%仍保持高准确率。
- 适合缺乏专业标注资源的音乐信息检索研究者使用。
监督学习依赖大规模标注数据集,但音频与音乐领域因标注成本高、需专业知识,缺乏充足标注数据。本文探索基于图的半监督学习方法——标签传播(LP),通过构建音频嵌入的相似性图,在归纳式设置下将少量标注数据传播至大规模无标签数据。我们在印度艺术音乐(IAM)的两个任务中应用该方法:拉格识别与乐器分类,整合多个公开数据集及来自普沙尔·巴哈里档案馆的新增录音。实验表明,相比传统基线方法(包括基于预训练归纳模型的方法),该方法显著降低标注开销,生成更高质量的标注结果。这证明了图基半监督学习在降低数据标注门槛、推动音乐信息检索发展方面的潜力。
原文摘要 · Abstract (English)
Supervised machine learning frameworks rely on extensive labeled datasets for robust performance on real-world tasks. However, there is a lack of large annotated datasets in audio and music domains, as annotating such recordings is resource-intensive, laborious, and often require expert domain knowledge. In this work, we explore the use of label propagation (LP), a graph-based semi-supervised learning technique, for automatically labeling the unlabeled set in an unsupervised manner. By constructing a similarity graph over audio embeddings, we propagate limited label information from a small annotated subset to a larger unlabeled corpus in a transductive, semi-supervised setting. We apply this method to two tasks in Indian Art Music (IAM): Raga identification and Instrument classification. For both these tasks, we integrate multiple public datasets along with additional recordings we acquire from Prasar Bharati Archives to perform LP. Our experiments demonstrate that LP significantly reduces labeling overhead and produces higher-quality annotations compared to conventional baseline methods, including those based on pretrained inductive models. These results highlight the potential of graph-based semi-supervised learning to democratize data annotation and accelerate progress in music information retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。