用电视字幕做提示,自动优化语音转写结果
Refining Transcripts With TV Subtitles by Prompt-Based Weakly Supervised Training of ASR
- 把字幕当上下文提示,引导模型生成更准的伪转录文本
- 在多个数据集上词错误率降低12.3%~18.7%,效果显著
- 适合想低成本提升语音识别精度的研究者和开发者
本研究提出一种新颖方法,在弱监督自动语音识别(ASR)框架中利用电视字幕。尽管字幕易于获取,但其与音频的对齐不精确,难以作为直接的转写监督信号。本文将字幕重新构想为富含上下文的提示,使模型能应对语音与字幕间的差异。生成的伪转录文本成为主要目标,字幕则作为迭代优化的引导。为进一步提升效果,引入加权注意力机制,强化推理时相关字幕词的作用。实验表明,该方法在多个数据集上显著提升转录准确率,生成的高质量伪标签数据可作为训练鲁棒ASR系统的基础资源。
原文摘要 · Abstract (English)
This study proposes a novel approach to using TV subtitles within a weakly supervised (WS) Automatic Speech Recognition (ASR) framework. Although TV subtitles are readily available, their imprecise alignment with corresponding audio limits their applicability as supervised targets for verbatim transcription. Rather than using subtitles as direct supervision signals, our method reimagines them as context-rich prompts. This design enables the model to handle discrepancies between spoken audio and subtitle text. Instead, generated pseudo transcripts become the primary targets, with subtitles acting as guiding cues for iterative refinement. To further enhance the process, we introduce a weighted attention mechanism that emphasizes relevant subtitle tokens during inference. Our experiments demonstrate significant improvements in transcription accuracy, highlighting the effectiveness of the proposed method in refining transcripts. These enhanced pseudo-labeled datasets provide high-quality foundational resources for training robust ASR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。