arXiv:2511.20006eess.AScs.AI2025-11中稿 · publication in IEE…

用音乐语言模型实现无参考音高矫正,保留演唱情感表达。

BERT-APC: A Reference-free Framework for Automatic Pitch Correction via Musical Context Inference

  • 通过音乐上下文推理预测目标音高,无需参考音轨。
  • 在高度走调样本中准确率比最优基线高10.49个百分点。
  • 适合音乐制作人和歌手,兼顾音准与表演自然度。

自动音高矫正(APC)通过将演唱中的音高偏差对齐到预期音符来提升录音质量。现有方法要么依赖参考音高(限制实用性),要么使用简单音高估计算法,常导致表现力丧失。本文提出BERT-APC,一种无参考的APC框架,在保持演唱表现力和自然性的同时纠正音高误差。首先,静态音高预测器从走调歌声中估计每个音符的稳定音高(即感知音高)。接着,基于重构的音乐语言模型,上下文感知音高预测器推断出意图的音高序列。最后,音符级校正算法修正音高错误,同时保留有意的情感偏差。我们还引入可学习的数据增强策略,模拟真实走调模式以提升鲁棒性。相较于两个近期声乐转录模型,BERT-APC在高度走调样本中原始音高准确率领先第二名ROSVOT达10.49个百分点。在主观听感测试(MOS)中,其得分高达4.32±0.15,显著优于Auto-Tune(3.22±0.18)和Melodyne(3.08±0.18),且保留表达细节能力相当。据我们所知,这是首个利用音乐语言模型实现符号化音乐上下文驱动的无参考音高矫正模型。修正音频示例见:https://joshua-1995.github.io/BERT-APC-Demo/

原文摘要 · Abstract (English)

Automatic Pitch Correction (APC) enhances vocal recordings by aligning pitch deviations with intended musical notes. However, existing APC systems either rely on reference pitches, which limits practical applicability, or employ simple pitch estimation algorithms that often fail to preserve expressiveness and naturalness. We propose BERT-APC, a reference-free APC framework that corrects pitch errors while maintaining the expressiveness and naturalness of vocal performances. In BERT-APC, a stationary pitch predictor first estimates the stationary pitch of each note from the detuned singing voice, where stationary pitch is the continuous pitch from the stable region of a note and approximates its perceived pitch. A context-aware note pitch predictor then infers the intended pitch sequence using a repurposed music language model that incorporates musical context. Finally, a note-level correction algorithm fixes pitch errors while preserving intentional deviations for emotional expression. We also introduce a learnable data augmentation strategy that improves robustness by simulating realistic detuning patterns. Compared to two recent singing voice transcription models, BERT-APC demonstrated superior target note pitch prediction, outperforming the second-best model, ROSVOT, by 10.49 percentage points on highly detuned samples in raw pitch accuracy. In the MOS test, BERT-APC achieved the highest quality rating of $4.32 \pm 0.15$, significantly higher than Auto-Tune ($3.22 \pm 0.18$) and Melodyne ($3.08 \pm 0.18$), while maintaining a comparable ability to preserve expressive nuances. To the best of our knowledge, this is the first APC model that leverages a music language model to achieve reference-free pitch correction with symbolic musical context. The corrected audio samples are available at https://joshua-1995.github.io/BERT-APC-Demo/.

音高矫正音乐语言模型无参考表现力保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。