arXiv:2601.12153eess.AScs.SD2026-01综述被引 2

综述30年歌唱评估技术,揭示标准缺失与艺术表现捕捉难题

A Survey on 30+ Years of Automatic Singing Assessment and Singing Information Processing

  • 梳理三十年歌唱评估技术演进路径
  • 指出缺乏统一评价框架与噪声分离难
  • 适合音乐科技、声学分析研究者参考

自动歌唱评估与歌唱信息处理在过去三十年中不断发展,以支持声乐教学、表演分析和发声训练。早期方法通过计算指标客观评估演唱表现,涵盖实时视觉反馈、声学生物反馈、精确的音高追踪与频谱分析;后期方法则通过比较预测声信号与目标参考信号,捕捉歌声中的细微特征。显著进展包括交互式系统提升实时反馈效果,以及机器学习与深度神经网络架构增强语音信号处理精度。本综述批判性分析相关文献,梳理技术发展脉络,识别并讨论关键空白。分析显示,持久挑战包括缺乏标准化评估框架、难以可靠分离语音信号与各类噪声源,以及在捕捉艺术表现力方面对先进数字信号处理与人工智能方法的利用不足。通过详述这些局限与对应技术进步,本文展示如何弥合客观计算评估与主观人类评价之间的差距,从而提升自动化歌唱评估系统的技术准确性和教学相关性。

原文摘要 · Abstract (English)

Automatic Singing Assessment and Singing Information Processing have evolved over the past three decades to support singing pedagogy, performance analysis, and vocal training. While the first approach objectively evaluates a singer's performance through computational metrics ranging from real-time visual feedback and acoustical biofeedback to sophisticated pitch tracking and spectral analysis, the latter method compares a predictor vocal signal with a target reference to capture nuanced data embedded in the singing voice. Notable advancements include the development of interactive systems that have significantly improved real-time visual feedback, and the integration of machine learning and deep neural network architectures that enhance the precision of vocal signal processing. This survey critically examines the literature to map the historical evolution of these technologies, while identifying and discussing key gaps. The analysis reveals persistent challenges, such as the lack of standardized evaluation frameworks, difficulties in reliably separating vocal signals from various noise sources, and the underutilization of advanced digital signal processing and artificial intelligence methodologies for capturing artistic expressivity. By detailing these limitations and the corresponding technological advances, this review demonstrates how addressing these issues can bridge the gap between objective computational assessments and subjective human-like evaluations of singing performance, ultimately enhancing both the technical accuracy and pedagogical relevance of automated singing evaluation systems.

歌唱评估声学分析人工智能音乐科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。