构建高质量越南语语音识别数据集,解决低资源语言数据不足问题。
Vietnamese Automatic Speech Recognition: A Revisit
- 从多种开源来源整合并清洗数据,自动标注词级时间戳。
- 生成500小时统一高质量数据集,显著提升模型训练效果。
- 方法通用可复用,适合低资源语言语音识别研究者使用。
自动语音识别(ASR)性能高度依赖大规模、高质量的数据集。对于低资源语言,现有开源数据集常因质量不足和标注不一致而限制模型发展。为此,我们提出一种新颖且可泛化的数据聚合与预处理流程,旨在从多样且可能嘈杂的开源源中构建高质量的ASR数据集。该流程包含严格的处理步骤,确保数据多样性、平衡性,并保留词级时间戳等关键信息。我们将该方法应用于越南语,成功构建了一个统一的500小时高质量数据集,为训练和评估最先进的越南语ASR系统提供了基础。项目主页见:https://github.com/qualcomm-ai-research/PhoASR。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) performance is heavily dependent on the availability of large-scale, high-quality datasets. For low-resource languages, existing open-source ASR datasets often suffer from insufficient quality and inconsistent annotation, hindering the development of robust models. To address these challenges, we propose a novel and generalizable data aggregation and preprocessing pipeline designed to construct high-quality ASR datasets from diverse, potentially noisy, open-source sources. Our pipeline incorporates rigorous processing steps to ensure data diversity, balance, and the inclusion of crucial features like word-level timestamps. We demonstrate the effectiveness of our methodology by applying it to Vietnamese, resulting in a unified, high-quality 500-hour dataset that provides a foundation for training and evaluating state-of-the-art Vietnamese ASR systems. Our project page is available at https://github.com/qualcomm-ai-research/PhoASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。