arXiv:2609.02901cs.CLcs.SD2026-09

让语音识别同时输出口语和可读文本,提升数字表达准确性。

Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition

论文配图:Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition
图 1 · 摘自论文原文
  • 用双形式监督训练模型,让语音识别直接生成语义感知的书面文本。
  • 在中文语音数据上,比开源系统更准,对数字等敏感内容纠错能力强。
  • 支持按需切换口语或书面输出,适合需要精确表达的场景。

现代语音识别需同时提供忠实转录的口语形式和可读的书面形式,后者依赖逆文本归一化(ITN)。传统方法将两者分步处理,导致书面输出易受识别错误影响,且归一化与声学上下文脱节,尤其在语义相关的数字表达中表现不佳。本文提出双形式语音识别(DF-ASR),通过配对的口语与书面形式监督,扩展语音识别能力以实现语义感知的书面归一化,同时保留输出形式的提示级选择。监督信号由大语言模型驱动的生成与判断流程构建,训练引入基于序列的ITN-MWER目标,对归一化敏感段落赋予更高惩罚。还设计了决策感知的REQUIRE-ITN/ FORBID-ITN协议,分别评估必需归一化与禁止性片段的保持能力。在SpeechIO中文标注子集上,DF-ASR持续优于开源的ASR-ITN系统,性能接近闭源强参考模型,并保持可靠的提示级形式控制。

原文摘要 · Abstract (English)

Modern automatic speech recognition (ASR) scenarios require both spoken-form transcripts for faithful transcription and readable written-form transcripts with inverse text normalization (ITN). However, these forms are typically produced by cascaded modules, where a spoken-form ASR output is rewritten by a separate ITN component, making written-form ASR-ITN vulnerable to recognition errors and decoupling normalization from acoustic-contextual modeling, especially for semantically dependent numeric expressions. In this paper, we propose Dual-Form ASR (DF-ASR), a framework that extends spoken-form ASR capability to semantics-aware written-form ITN through paired spoken-form and written-form supervision while retaining prompt-level selection between transcript forms. The dual-form supervision is constructed via a large language model (LLM)-driven generate-and-judge workflow, and training is further enhanced by ITN-MWER, a sequence-level objective that assigns higher cost to errors on normalization-sensitive spans. We also introduce a decision-aware REQUIRE-ITN/\FORBID-ITN protocol to separately measure required normalization and forbidden-span preservation. On manually annotated Chinese subsets from SpeechIO, DF-ASR consistently outperforms open-source ASR-ITN systems, remains competitive with strong closed-source references, and preserves reliable prompt-level control between spoken-form and written-form outputs.

语音识别文本归一化中文处理双形式输出

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。