arXiv:2505.11378cs.SDcs.LG2025-05被引 1

用音频特征自动识别男声流行唱法的发声类型

Machine Learning Approaches to Vocal Register Classification in Contemporary Male Pop Music

  • 通过梅尔频谱图纹理特征分析,区分男声在流行音乐中的发声状态
  • SVM与CNN模型均实现稳定分类,准确率表现可靠
  • 适用于声乐教学工具,适合音乐教育者与歌手使用

对于各水平的歌手而言,掌握发声位置和音域转换区(即胸声与头声之间的过渡区)是学习技巧曲目的重大挑战。尤其在流行音乐中,同一歌手可能频繁切换不同音色与质感以达到理想表现,使判断其实际使用的发声类型变得困难。本文提出两种基于音频信号中梅尔频谱图纹理特征的男声发声类型分类方法。我们还探讨了这些模型在声乐分析工具中的实际应用,并介绍了同步开发的软件AVRA(Automatic Vocal Register Analysis)。所提方法在支持向量机(SVM)与卷积神经网络(CNN)模型上均实现了稳定的发声类型分类,表明未来有望扩展至更多嗓音类型与音乐风格。

原文摘要 · Abstract (English)

For singers of all experience levels, one of the most daunting challenges in learning technical repertoire is navigating placement and vocal register in and around the passagio (passage between chest voice and head voice registers). Particularly in pop music, where a single artist may use a variety of timbre's and textures to achieve a desired quality, it can be difficult to identify what vocal register within the vocal range a singer is using. This paper presents two methods for classifying vocal registers in an audio signal of male pop music through the analysis of textural features of mel-spectrogram images. Additionally, we will discuss the practical integration of these models for vocal analysis tools, and introduce a concurrently developed software called AVRA which stands for Automatic Vocal Register Analysis. Our proposed methods achieved consistent classification of vocal register through both Support Vector Machine (SVM) and Convolutional Neural Network (CNN) models, which supports the promise of more robust classification possibilities across more voice types and genres of singing.

声乐分析语音分类深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。