构建首个声乐模式自动分类数据集,助力智能声乐教学。
A Dataset for Automatic Vocal Mode Classification
- 采集4名歌手的持续元音,覆盖完整音域,共3752个独特样本
- 采用4麦克风录制,总样本超1.3万,标注由3位专家独立完成
- 基于ResNet18实现81.3%平衡准确率,可为声乐教学工具提供支持
声乐技术(CVT)将发声方式分为中性、抑制、过载和边缘四种声乐模式。掌握目标模式对声乐学习者有帮助,因此自动分类具有重要价值。此前研究因数据不足而进展有限。本文构建了首个针对该任务的声乐模式数据集,包含4名歌手(3位专业歌手,经验超5年)的持续元音录音,覆盖全部音域,共3,752个唯一样本。通过4麦克风采集,总样本数超过13,000。数据集附带三位CVT专家的独立标注及合并标注。同时提供基线结果:在5折交叉验证下,使用ResNet18模型达到81.3%的平衡准确率。数据集已发布于Zenodo。
原文摘要 · Abstract (English)
The Complete Vocal Technique (CVT) is a school of singing developed in the past decades by Cathrin Sadolin et al.. CVT groups the use of the voice into so called vocal modes, namely Neutral, Curbing, Overdrive and Edge. Knowledge of the desired vocal mode can be helpful for singing students. Automatic classification of vocal modes can thus be important for technology-assisted singing teaching. Previously, automatic classification of vocal modes has been attempted without major success, potentially due to a lack of data. Therefore, we recorded a novel vocal mode dataset consisting of sustained vowels recorded from four singers, three of which professional singers with more than five years of CVT-experience. The dataset covers the entire vocal range of the subjects, totaling 3,752 unique samples. By using four microphones, thereby offering a natural data augmentation, the dataset consists of more than 13,000 samples combined. An annotation was created using three CVT-experienced annotators, each providing an individual annotation. The merged annotation as well as the three individual annotations come with the published dataset. Additionally, we provide some baseline classification results. The best balanced accuracy across a 5-fold cross validation of 81.3\,\% was achieved with a ResNet18. The dataset can be downloaded under https://zenodo.org/records/14276415.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。