arXiv:2505.24229cs.CLcs.SD2025-05中稿 · INTERSPEECH 2025被引 1

提出动态上下文感知的流式ITN模型,提升低资源场景下的准确率与效率。

Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization

  • 训练与推理时动态调整分块大小,融合前后文信息以适应流式输入。
  • 在越南语数据集上达到非流式模型水平准确率,且延迟更低。
  • 适合语音识别系统中实时文本规范化,尤其适用于资源有限场景。

逆文本归一化(ITN)对于将语音识别(ASR)输出的口语化文本转化为规范书面语至关重要,可显著提升可读性与可用性。尽管其重要性突出,但将流式ITN集成到流式ASR中仍面临准确性、效率和适应性挑战,尤其是在低资源和短上下文场景下。本文提出一种面向流式ITN的预训练语言模型,利用预训练的语言表征提升鲁棒性。为应对流式限制,我们引入训练与推理阶段的动态上下文感知机制,实现自适应分块大小调整,并整合右向上下文信息。实验表明,该方法在越南语数据集上达到与非流式ITN相当的准确率,优于现有流式ITN模型,同时保持低延迟,可无缝集成至ASR系统。

原文摘要 · Abstract (English)

Inverse Text Normalization (ITN) is crucial for converting spoken Automatic Speech Recognition (ASR) outputs into well-formatted written text, enhancing both readability and usability. Despite its importance, the integration of streaming ITN within streaming ASR remains largely unexplored due to challenges in accuracy, efficiency, and adaptability, particularly in low-resource and limited-context scenarios. In this paper, we introduce a streaming pretrained language model for ITN, leveraging pretrained linguistic representations for improved robustness. To address streaming constraints, we propose Dynamic Context-Aware during training and inference, enabling adaptive chunk size adjustments and the integration of right-context information. Experimental results demonstrate that our method achieves accuracy comparable to non-streaming ITN and surpasses existing streaming ITN models on a Vietnamese dataset, all while maintaining low latency, ensuring seamless integration into ASR systems.

逆文本归一化流式处理语音识别预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。