用半监督和大模型优化爱沙尼亚语字幕,逼近人工水平。
Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs
- 基于伪标签迭代优化Whisper模型,提升未标注数据利用效率。
- 测试时用大模型后编辑使字幕准确率显著提升,训练时则无效。
- 适合需要高质量本地化字幕的影视平台,尤其资源有限的语种。
本文提出一种生成爱沙尼亚语电视内容高质量同语言字幕的方法。通过在人工生成的爱沙尼亚语字幕上微调Whisper模型,并结合迭代伪标签与大语言模型(LLM)的后编辑进行增强。实验表明,使用未标注数据进行伪标签迭代可显著提升字幕质量。研究发现,测试阶段应用LLM编辑能有效提高字幕准确性,而训练阶段使用则无额外增益。该方法有望实现接近人工标准的字幕质量,具备向实时应用扩展的潜力。
原文摘要 · Abstract (English)
This paper presents an approach for generating high-quality, same-language subtitles for Estonian TV content. We fine-tune the Whisper model on human-generated Estonian subtitles and enhance it with iterative pseudo-labeling and large language model (LLM) based post-editing. Our experiments demonstrate notable subtitle quality improvement through pseudo-labeling with an unlabeled dataset. We find that applying LLM-based editing at test time enhances subtitle accuracy, while its use during training does not yield further gains. This approach holds promise for creating subtitle quality close to human standard and could be extended to real-time applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。