用HerBERT模型提升波兰语语音转写文本的标点预测准确率
Punctuation Prediction for Polish Texts using Transformers
- 基于HerBERT模型微调,融合竞赛与外部数据
- 在Poleval 2022任务中达到71.44的加权F1分数
- 适合需要提升波兰语文本可读性的应用场景
语音识别系统通常输出无标点的文本。然而,标点对书面文本的理解至关重要。为解决此问题,本文针对Poleval 2022任务1:波兰语文本的标点预测,提出一种解决方案。该方法使用单个HerBERT模型,在竞赛数据与外部数据上进行微调,最终获得71.44的加权F1分数。该结果展示了基于预训练语言模型在波兰语标点恢复任务中的有效性。
原文摘要 · Abstract (English)
Speech recognition systems typically output text lacking punctuation. However, punctuation is crucial for written text comprehension. To tackle this problem, Punctuation Prediction models are developed. This paper describes a solution for Poleval 2022 Task 1: Punctuation Prediction for Polish Texts, which scores 71.44 Weighted F1. The method utilizes a single HerBERT model finetuned to the competition data and an external dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。