用电视字幕弱监督数据提升语音识别与自动字幕生成效果
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
- 将字幕与原声文本视为不同语域,用双编码器联合建模
- 在弗拉芒语上实现语音识别与字幕生成双重性能提升
- 无需复杂预处理,可扩展至大规模字幕数据
语音识别技术的进步依赖于大规模数据集和基于注意力的架构,但低资源语言和方言仍面临诸多挑战。本文探索将电视字幕的弱监督转录文本融入自动语音识别(ASR)系统,以同时提升逐字转录和自动生成字幕的效果。由于转录文本与字幕具有不同特征,二者被视为不同领域或语言。我们提出并对比多种端到端架构,通过独立或共享编码器与解码器联合建模两种模态。所提方法可同步生成逐字转录与字幕。在弗拉芒语(比利时荷兰语)上的实验表明,采用级联编码器与独立解码器的模型最有效捕捉两者差异,并在两个任务上均取得改进。尽管存在领域差异与语言变体,结合转录文本与字幕数据显著提升了ASR性能,且无需大量预处理。此外,使用大规模字幕数据集的实验验证了该方法的可扩展性。该方法不仅提高语音识别准确率,还生成与标准书面语高度匹配的字幕,具备广泛应用潜力。
原文摘要 · Abstract (English)
The recent advancement of speech recognition technology has been driven by large-scale datasets and attention-based architectures, but many challenges still remain, especially for low-resource languages and dialects. This paper explores the integration of weakly supervised transcripts from TV subtitles into automatic speech recognition (ASR) systems, aiming to improve both verbatim transcriptions and automatically generated subtitles. To this end, verbatim data and subtitles are regarded as different domains or languages, due to their distinct characteristics. We propose and compare several end-to-end architectures that are designed to jointly model both modalities with separate or shared encoders and decoders. The proposed methods are able to jointly generate a verbatim transcription and a subtitle. Evaluation on Flemish (Belgian Dutch) demonstrates that a model with cascaded encoders and separate decoders allows to represent the differences between the two data types most efficiently while improving on both domains. Despite differences in domain and linguistic variations, combining verbatim transcripts with subtitle data leads to notable ASR improvements without the need for extensive preprocessing. Additionally, experiments with a large-scale subtitle dataset show the scalability of the proposed approach. The methods not only improve ASR accuracy but also generate subtitles that closely match standard written text, offering several potential applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。