arXiv:2505.23170cs.CLcs.SD2025-05ACL被引 17

ZIPA模型用更少参数实现多语言语音识别新高度。

ZIPA: A family of efficient models for multilingual phone recognition

  • 基于高效Zipformer结构,构建跨语言语音识别模型
  • 在17,132小时数据上训练,参数更少但性能更强
  • 适合需要轻量级多语种语音识别的场景

我们提出ZIPA,一类高效的语音模型,显著提升跨语言音素识别的性能。首先构建了包含17,132小时标准化音素转录的大规模多语言语料库IPAPack++,并设计了一个涵盖未见语言和社会语音变化的新颖评估集。利用大规模训练数据,ZIPA(包括基于转换器的ZIPA-T和基于CTC的ZIPA-CR)采用高效的Zipformer骨干网络,在参数远少于现有系统的情况下仍表现更优。通过在11,000小时伪标签多语言数据上进行噪声学生训练进一步提升性能。尽管在基准测试中表现强劲,错误分析揭示其在建模社会语音多样性方面仍存在持续局限,凸显未来研究挑战。

原文摘要 · Abstract (English)

We present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. We first curated IPAPack++, a large-scale multilingual speech corpus with 17,132 hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation. With the large-scale training data, ZIPA, including transducer (ZIPA-T) and CTC-based (ZIPA-CR) variants, leverage the efficient Zipformer backbones and outperform existing phone recognition systems with much fewer parameters. Further scaling via noisy student training on 11,000 hours of pseudo-labeled multilingual data yields further improvement. While ZIPA achieves strong performance on benchmarks, error analysis reveals persistent limitations in modeling sociophonetic diversity, underscoring challenges for future research.

语音识别多语言轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。