arXiv:2602.20744cs.SDcs.AI2026-02

首个针对库尔德马卡姆音乐的声乐错误检测系统,突破西方音乐标准局限。

Voices of the Mountains: Deep Learning-Based Vocal Error Detection System for Kurdish Maqams

  • 采用双头CNN-BiLSTM+注意力模型分析微音程、节奏与调式稳定性误差
  • 在50首歌曲上实现39.4%召回率,节奏错误识别准确率达53.6%
  • 专为非平均律的库尔德传统音乐设计,适合民族音乐研究与教学应用

马卡姆是库尔德音乐的重要歌唱形式,其演唱依赖传统面授或自学。自动声乐评估(ASA)利用机器学习技术可帮助学习者通过错误检测提升表现。现有工具多遵循西方音乐规则,要求所有音符严格保持在固定音高范围内,无法识别微音程和音高滑动,导致库尔德马卡姆演唱被误判为错误。而库尔德马卡姆需在微音程空间内判断表演误差,超出西方十二平均律范畴。本研究首次填补该空白。聚焦巴亚蒂-库尔德调式中的音高、节奏与调式稳定性错误,收集13位歌手共50首歌曲(2-3小时),标注221个错误段落(细音高150处,节奏46处,调式漂移25处)。数据分割为15,199个重叠窗口并转换为对数梅尔频谱图。构建双头CNN-BiLSTM带注意力机制模型,用于判断窗口是否含错及分类。训练20轮,早停于第10轮,验证集宏F1达0.468。全量50首测试中,阈值0.75时召回率为39.4%,精确率为25.8%。在检测窗口内,类型宏F1为0.387,其中细音高F1为0.492,节奏为0.536,调式漂移仅为0.133,调式漂移召回率仅8.0%。模型在常见错误类型上表现良好,但调式漂移识别仍需更多数据与平衡优化。

原文摘要 · Abstract (English)

Maqam, a singing type, is a significant component of Kurdish music. A maqam singer receives training in a traditional face-to-face or through self-training. Automatic Singing Assessment (ASA) uses machine learning (ML) to provide the accuracy of singing styles and can help learners to improve their performance through error detection. Currently, the available ASA tools follow Western music rules. The musical composition requires all notes to stay within their expected pitch range from start to finish. The system fails to detect micro-intervals and pitch bends, so it identifies Kurdish maqam singing as incorrect even though the singer performs according to traditional rules. Kurdish maqam requires recognizing performance errors within microtonal spaces, which is beyond Western equal temperament. This research is the first attempt to address the mentioned gap. While many error types happen during singing, our focus is on pitch, rhythm, and modal stability errors in the context of Bayati-Kurd. We collected 50 songs from 13 vocalists ( 2-3 hours) and annotated 221 error spans (150 fine pitch, 46 rhythm, 25 modal drift). The data was segmented into 15,199 overlapping windows and converted to log-mel spectrograms. We developed a two-headed CNN-BiLSTM with attention mode to decide whether a window contains an error and to classify it based on the chosen errors. Trained for 20 epochs with early stopping at epoch 10, the model reached a validation macro-F1 of 0.468. On the full 50-song evaluation at a 0.750 threshold, recall was 39.4% and precision 25.8% . Within detected windows, type macro-F1 was 0.387, with F1 of 0.492 (fine pitch), 0.536 (rhythm), and 0.133 (modal drift); modal drift recall was 8.0%. The better performance on common error types shows that the method works, while the poor modal-drift recall shows that more data and balancing are needed.

声乐评估微音程民族音乐深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。