arXiv:2603.19259cs.CLcs.AI2026-03

构建台湾话语音识别与合成的标准化评测框架

Breeze Taigi: Benchmarks and Models for Taiwanese Hokkien Speech Recognition and Synthesis

  • 基于台语与国语平行语料,建立可复现的评测方法
  • 用10,000小时合成数据微调Whisper模型,平均字错误率30.13%
  • 提供开源基准模型,适合多语言语音技术研究者使用

台湾话(Taigi)为推动语音技术方法在多样化语言环境中的泛化提供了独特机遇。本文提出Breeze Taigi,一个以标准化评测为核心的综合性框架,用于评估台湾话语音识别与合成系统。核心贡献在于可复现的评估方法,利用平行的台湾国语资源。我们提供来自台湾行政院公共服务公告的30组精心筛选的国语-台语音频对,附带标准化转写文本。确立字符错误率(CER)为标准指标,并实施归一化流程,确保跨系统公平比较。为展示基准实用性并提供参考实现,我们通过利用现有台湾国语资源和大规模合成数据生成的方法,开发了语音识别与合成模型。特别地,我们在约10,000小时的台语合成语音数据上微调Whisper模型。该语音识别模型在基准测试中达到平均30.13%的字符错误率,优于现有商业与研究系统。通过提供标准化评估协议、多样训练数据集及开放基线模型,我们构建了一个可复制的框架,其方法适用于多种语言场景。

原文摘要 · Abstract (English)

Taiwanese Hokkien (Taigi) presents unique opportunities for advancing speech technology methodologies that can generalize to diverse linguistic contexts. We introduce Breeze Taigi, a comprehensive framework centered on standardized benchmarks for evaluating Taigi speech recognition and synthesis systems. Our primary contribution is a reproducible evaluation methodology that leverages parallel Taiwanese Mandarin resources. We provide 30 carefully curated Mandarin-Taigi audio pairs from Taiwan's Executive Yuan public service announcements with normalized ground truth transcriptions. We establish Character Error Rate (CER) as the standard metric and implement normalization procedures to enable fair cross-system comparisons. To demonstrate the benchmark's utility and provide reference implementations, we develop speech recognition and synthesis models through a methodology that leverages existing Taiwanese Mandarin resources and large-scale synthetic data generation. In particular, we fine-tune a Whisper model on approximately 10,000 hours of Taigi synthetic speech data. Our ASR model achieves 30.13% average CER on the benchmark, outperforming existing commercial and research systems. By providing standardized evaluation protocols, diverse training datasets, and open baseline models, we offer a replicable framework with methodologies applicable to various linguistic contexts.

语音识别台语合成数据Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。