用wav2vec2自动估算语音起始时间等发音特征,准确率高且可泛化。
wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2
- 基于wav2vec2模型,端到端估算声门起始时间、闭塞时长和爆破特征。
- 在未见数据集上表现接近现有方法,微调后精度显著提升。
- 适合语音学研究者,尤其关注自动化标注与大规模语料分析者。
尽管语音标注的自动化工具已在语音学研究中普及,但许多任务仍需大量人工校正或训练数据才能达到准确度。与此同时,像wav2vec2这样的大型语音模型在语音分类任务中表现出色,这引发了其在语音学标注任务中应用的可能性。本文提出wav2VOT:一个利用wav2vec2自动估算声门起始时间(Voice Onset Time)、闭塞时长(Closure Duration)和爆破实现(Burst Realisation)的工具。实验表明,wav2VOT在未见数据集上的表现可媲美当前主流方法,经微调后能实现高精度估计。对预测结果的分析显示,该模型在不同塞音的清浊对立及发音部位上均具有高度保真性。这些结果证明大型语音模型具备生成精准语音标注的能力,进一步推动其在语音学研究流程中的探索与应用。
原文摘要 · Abstract (English)
While automatic tools for speech annotation are now commonplace within phonetic research pipelines, many tasks require substantial manual correction or training sets to perform accurately. Simultaneously, large speech models such as wav2vec2 have been shown to perform well at speech classification tasks, raising the question of how these models may be applied to phonetic annotation tasks. We introduce wav2VOT: a tool for the automatic estimation of voice onset time, closure duration, and burst realisation using wav2vec2. We demonstrate that wav2VOT performs comparably with current approaches on unseen datasets, and can estimate with high accuracy with fine-tuning. Analysis of wav2VOT predictions demonstrate high fidelity across stop voicing and place of articulation. These results demonstrate that large speech models are capable of producing accurate annotations, and further motivate exploration of large speech models as tools in phonetic research pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。