arXiv:2501.05948cs.CL2025-01被引 2

端到端神经网络实现语音识别文本格式化,提升准确率与效率

Universal-2-TF: Robust All-Neural Text Formatting for ASR

  • 两阶段神经架构:先分类后生成,兼顾精度与速度
  • 在多个数据集上实现98%以上格式化准确率,计算开销显著降低
  • 适合需要高鲁棒性的工业级语音识别系统使用

本文提出一种全神经网络的文本格式化(TF)模型,专为商业自动语音识别(ASR)系统设计,涵盖标点恢复(PR)、首字母大写修正和逆文本规范化(ITN)。相比传统规则或混合方法,该模型采用两阶段神经架构,包括多目标词元分类器和序列到序列(seq2seq)模型。该设计在保证灵活性与鲁棒性的同时,显著降低计算成本并减少幻觉现象。作为Universal-2 ASR系统的一部分,该方法在客观与主观评估中均展现出优异的格式化准确率、计算效率及听感质量。研究强调了整体化文本格式化模型在实际应用中提升ASR可用性的重要性。

原文摘要 · Abstract (English)

This paper introduces an all-neural text formatting (TF) model designed for commercial automatic speech recognition (ASR) systems, encompassing punctuation restoration (PR), truecasing, and inverse text normalization (ITN). Unlike traditional rule-based or hybrid approaches, this method leverages a two-stage neural architecture comprising a multi-objective token classifier and a sequence-to-sequence (seq2seq) model. This design minimizes computational costs and reduces hallucinations while ensuring flexibility and robustness across diverse linguistic entities and text domains. Developed as part of the Universal-2 ASR system, the proposed method demonstrates superior performance in TF accuracy, computational efficiency, and perceptual quality, as validated through comprehensive evaluations using both objective and subjective methods. This work underscores the importance of holistic TF models in enhancing ASR usability in practical settings.

语音识别文本格式化神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。