arXiv:2409.09646cs.CL2024-09中稿 · SLT 2024被引 2

用梅尔频谱峰检测做基线,结合自监督表示提升语音分割性能。

A Simple HMM with Self-Supervised Representations for Phone Segmentation

  • 基于梅尔频谱峰检测构建简单隐马尔可夫模型
  • 在多个数据集上优于主流自监督方法,提升明显
  • 适合语音识别、声学建模等语音处理任务研究者

尽管自监督表示近期取得进展,无监督音素分割仍具挑战性。多数方法致力于通过自监督学习改进音素表示,期望其能迁移至音素分割任务。本文相反地表明,对梅尔频谱图进行峰值检测是一种强基线,优于许多自监督方法。基于此发现,我们提出一种简单的隐马尔可夫模型,利用自监督表示及边界特征进行音素分割。实验结果表明,该方法在多个基准上持续优于先前方法,且具备通用形式,支持灵活设计调整。

原文摘要 · Abstract (English)

Despite the recent advance in self-supervised representations, unsupervised phonetic segmentation remains challenging. Most approaches focus on improving phonetic representations with self-supervised learning, with the hope that the improvement can transfer to phonetic segmentation. In this paper, contrary to recent approaches, we show that peak detection on Mel spectrograms is a strong baseline, better than many self-supervised approaches. Based on this finding, we propose a simple hidden Markov model that uses self-supervised representations and features at the boundaries for phone segmentation. Our results demonstrate consistent improvements over previous approaches, with a generalized formulation allowing versatile design adaptations.

语音分割自监督隐马尔可夫模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。