arXiv:2512.18371eess.AScs.SD2025-12被引 3

用随机采样提升语音识别模型训练效率与准确率

Phoneme-based speech recognition driven by large language models and sampling marginalization

  • 用随机采样替代束搜索生成候选音素路径
  • 在波兰语和德语数据集上加速收敛并提升识别精度
  • 适合追求高效低资源的跨语言语音识别系统

基于大语言模型的音素到音节映射(LLM-P2G)方法在语音识别中表现优异,已成为替代传统WFST解码的可行方向。该框架通过音素预测与文本生成两阶段建模,在提升识别准确率的同时增强系统可扩展性。然而现有方法采用基于Top-K的边际化训练策略(TKM),其候选音素序列依赖束搜索生成,存在路径多样性不足、训练效率低及资源开销高的问题。为此,本文提出采样边际化训练策略(Sampling-K Marginalized, SKM),以随机采样替代束搜索生成候选路径,改进边际化建模并提升训练效率。在波兰语和德语数据集上的实验表明,SKM显著加快模型学习收敛速度并提升识别性能,同时保持模型复杂度不变。与结合投影器的大语言模型语音识别方法(SpeechLLM)相比,SKM驱动的LLM-P2G在识别准确率和结构简洁性上均更具优势。研究验证了该方法在跨语言语音识别系统中的实用价值与应用潜力。

原文摘要 · Abstract (English)

Recently, the Large Language Model-based Phoneme-to-Grapheme (LLM-P2G) method has shown excellent performance in speech recognition tasks and has become a feasible direction to replace the traditional WFST decoding method. This framework takes into account both recognition accuracy and system scalability through two-stage modeling of phoneme prediction and text generation. However, the existing LLM-P2G adopts the Top-K Marginalized (TKM) training strategy, and its candidate phoneme sequences rely on beam search generation, which has problems such as insufficient path diversity, low training efficiency, and high resource overhead. To this end, this paper proposes a sampling marginalized training strategy (Sampling-K Marginalized, SKM), which replaces beam search with random sampling to generate candidate paths, improving marginalized modeling and training efficiency. Experiments were conducted on Polish and German datasets, and the results showed that SKM further improved the model learning convergence speed and recognition performance while maintaining the complexity of the model. Comparative experiments with a speech recognition method that uses a projector combined with a large language model (SpeechLLM) also show that the SKM-driven LLM-P2G has more advantages in recognition accuracy and structural simplicity. The study verified the practical value and application potential of this method in cross-language speech recognition systems.

语音识别大模型采样优化音素建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。