无需标注数据,轻量级模型自动精准提取音乐基频与发音状态。
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
- 基于时频特征的自监督学习,通过一致性信号筛选有效音帧。
- 在MedleyDB上实现95.84%的音高准确率和96.24%的发音准确率。
- 适合资源有限场景下的单乐器快速训练,无需人工标注。
可靠的基频(F₀)与发音状态估计对神经音频合成至关重要,但许多现有方法依赖大规模标注数据,在真实录音噪声下性能下降。本文提出一种轻量级、完全自监督的联合F₀估计与发音推断框架,适用于从少量音频中快速训练单乐器模型。利用对转调不变的CQT特征,引入类EM迭代加权机制,以移位交叉熵(SCE)一致性作为可靠性信号,抑制无信息的噪声或非发音帧。生成的权重提供置信度分数,用于为独立轻量级发音分类器进行伪标注,无需人工标注。在MedleyDB上训练,于MDB-stem-synth真实数据上评估,达到95.84%的跨语料音高准确率(RPA)和96.24%的发音准确率(RCA),并展现出跨乐器泛化能力。
原文摘要 · Abstract (English)
Reliable fundamental frequency (F 0) and voicing estimation is essential for neural synthesis, yet many pitch extractors depend on large labeled corpora and degrade under realistic recording artifacts. We propose a lightweight, fully self-supervised framework for joint F 0 estimation and voicing inference, designed for rapid single-instrument training from limited audio. Using transposition-equivariant learning on CQT features, we introduce an EM-style iterative reweighting scheme that uses Shift Cross-Entropy (SCE) consistency as a reliability signal to suppress uninformative noisy/unvoiced frames. The resulting weights provide confidence scores that enable pseudo-labeling for a separate lightweight voicing classifier without manual annotations. Trained on MedleyDB and evaluated on MDB-stem-synth ground truth, our method achieves competitive cross-corpus performance (RPA 95.84, RCA 96.24) and demonstrates cross-instrument generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。