arXiv:2509.04507cs.CLcs.AI2025-09被引 1

用双阶段模型提升无声语音的可懂度,让机器更准确听懂。

From Silent Signals to Natural Language: A Dual-Stage Transformer-LLM Approach

  • 先用变压器捕捉完整语句上下文,再用大语言模型校正语义
  • 字错误率降低16%相对、6%绝对,比基线提升36%
  • 适合需要高精度识别无声语音的场景,如医疗或军事应用

无声语音接口(SSIs)能够从非声学信号生成可理解语音。尽管语音生成流程已取得进展,但对合成语音的识别与后续处理仍面临音素模糊和噪声问题。为此,我们提出一种结合基于变压器的声学模型与大语言模型(LLM)的增强型自动语音识别框架。变压器捕获整句上下文,而LLM确保语言一致性。实验表明,相比36%的基线,字错误率(WER)实现16%相对与6%绝对降低,显著提升了无声语音接口的可懂度。

原文摘要 · Abstract (English)

Silent Speech Interfaces (SSIs) have gained attention for their ability to generate intelligible speech from non-acoustic signals. While significant progress has been made in advancing speech generation pipelines, limited work has addressed the recognition and downstream processing of synthesized speech, which often suffers from phonetic ambiguity and noise. To overcome these challenges, we propose an enhanced automatic speech recognition framework that combines a transformer-based acoustic model with a large language model (LLM) for post-processing. The transformer captures full utterance context, while the LLM ensures linguistic consistency. Experimental results show a 16% relative and 6% absolute reduction in word error rate (WER) over a 36% baseline, demonstrating substantial improvements in intelligibility for silent speech interfaces.

无声语音语音识别大模型双阶段

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。