arXiv:2508.09868cs.SD2025-08中稿 · presentation at IE…被引 1

对比不同语音识别架构在领域偏移下的表现,发现建模细节比模型类型更重要。

Analysis of Domain Shift across ASR Architectures via TTS-Enabled Separation of Target Domain and Acoustic Conditions

  • 用语音合成生成目标领域音频,分离语言域与声学变化影响
  • 在领域偏移下,特定建模选择比模型架构更影响识别性能
  • 无需重训练声学模型即可通过语言模型实现领域自适应

我们分析了自动语音识别(ASR)在领域不匹配条件下的建模选择,对比了经典模块化与新型序列到序列(seq2seq)架构。研究覆盖标签单元、上下文长度和网络拓扑等建模选择。为分离语言域效应与声学变化,我们使用基于LibriSpeech训练的文本到语音系统合成目标域音频。引入目标域n元语法和神经语言模型进行领域自适应,无需重训练声学模型。据我们所知,这是首个在领域偏移下对优化后的先进ASR系统进行受控比较的研究,揭示了其泛化能力。结果表明,在领域偏移条件下,影响性能的并非解码器架构或传统模块化与seq2seq的区别,而是具体的建模选择。

原文摘要 · Abstract (English)

We analyze automatic speech recognition (ASR) modeling choices under domain mismatch, comparing classic modular and novel sequence-to-sequence (seq2seq) architectures. Across the different ASR architectures, we examine a spectrum of modeling choices, including label units, context length, and topology. To isolate language domain effects from acoustic variation, we synthesize target domain audio using a text-to-speech system trained on LibriSpeech. We incorporate target domain n-gram and neural language models for domain adaptation without retraining the acoustic model. To our knowledge, this is the first controlled comparison of optimized ASR systems across state-of-the-art architectures under domain shift, offering insights into their generalization. The results show that, under domain shift, rather than the decoder architecture choice or the distinction between classic modular and novel seq2seq models, it is specific modeling choices that influence performance.

语音识别领域偏移语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。