arXiv:2503.10086cs.SDcs.MM2025-03中稿 · ISMIR2024被引 1

用自监督特征和适配器微调,提升人声节奏与强拍联合追踪效果

Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features

  • 结合自监督语音特征与频谱特征进行融合
  • 适配器微调使性能提升31.6%(节拍)和42.4%(强拍)
  • 特别适合处理非均质人声数据,适合音频分析研究者

人声节拍追踪因缺乏伴奏中稳定的节奏与和声模式而极具挑战性,现有系统多依赖伴奏中的强节奏线索。本文提出一种基于时序卷积网络的联合节拍与强拍追踪方法,利用自监督学习(SSL)的DistilHuBERT特征捕捉人声语义信息,并与通用频谱特征融合以增强节拍估计。通过高效的适配器微调策略,有效降低非均质人声数据带来的差异性影响。大量实验表明,特征融合与适配器微调均独立提升性能,二者结合相较未适配基线系统,在节拍追踪上实现最高31.6%的绝对F1分数提升,强拍追踪提升达42.4%。

原文摘要 · Abstract (English)

Singing voice beat tracking is a challenging task, due to the lack of musical accompaniment that often contains robust rhythmic and harmonic patterns, something most existing beat tracking systems utilize and can be essential for estimating beats. In this paper, a novel temporal convolutional network-based beat-tracking approach featuring self-supervised learning (SSL) representations and adapter tuning is proposed to track the beat and downbeat of singing voices jointly. The SSL DistilHuBERT representations are utilized to capture the semantic information of singing voices and are further fused with the generic spectral features to facilitate beat estimation. Sources of variabilities that are particularly prominent with the non-homogeneous singing voice data are reduced by the efficient adapter tuning. Extensive experiments show that feature fusion and adapter tuning improve the performance individually, and the combination of both leads to significantly better performances than the un-adapted baseline system, with up to 31.6% and 42.4% absolute F1-score improvements on beat and downbeat tracking, respectively.

节拍追踪自监督学习适配器微调人声分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。