构建40小时牙买加帕多瓦斯音乐语料库,提升语音识别准确率。
Towards Robust Speech Recognition for Jamaican Patois Music Transcription
- 收集40小时人工标注的牙买加帕多瓦斯音乐数据,用于模型微调。
- 基于该数据集建立Whisper模型在帕多瓦斯语音上的性能扩展规律。
- 为方言语音识别和文化内容可及性提供新方法,适合语言技术研究者。
尽管牙买加帕多瓦斯语使用广泛,现有语音识别系统在处理帕多瓦斯音乐时表现不佳,生成不准确的字幕,限制了可访问性并阻碍下游应用。本文采用数据驱动方法,整理超过40小时的人工标注帕多瓦斯音乐数据,以此微调先进的自动语音识别(ASR)模型,并基于结果推导出Whisper模型在牙买加帕多瓦斯语音上的性能扩展规律。本工作有望提升帕多瓦斯音乐的可访问性,并推动该语言建模的未来发展。
原文摘要 · Abstract (English)
Although Jamaican Patois is a widely spoken language, current speech recognition systems perform poorly on Patois music, producing inaccurate captions that limit accessibility and hinder downstream applications. In this work, we take a data-centric approach to this problem by curating more than 40 hours of manually transcribed Patois music. We use this dataset to fine-tune state-of-the-art automatic speech recognition (ASR) models, and use the results to develop scaling laws for the performance of Whisper models on Jamaican Patois audio. We hope that this work will have a positive impact on the accessibility of Jamaican Patois music and the future of Jamaican Patois language modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。