arXiv:2507.18448cs.CLcs.AI2025-07被引 3

用Transformer模型提升孟加拉语无标点文本的可读性

Restoring Rhythm: Punctuation Restoration Using Transformer Models for Bangla, A Low-Resource Language

  • 采用XLM-RoBERTa-large模型预测四种标点符号
  • 在新闻数据集上达97.1%准确率,支持真实语音识别场景
  • 公开数据集与代码,助力低资源语言NLP研究

标点恢复能提升文本可读性,对自动语音识别(ASR)等后处理任务至关重要,尤其适用于孟加拉语这类低资源语言。本文探索使用基于Transformer的XLM-RoBERTa-large模型,自动恢复无标点孟加拉语文本中的句号、逗号、问号和感叹号。针对标注资源稀缺问题,构建了大规模多样化训练语料,并应用数据增强技术。最佳模型在新闻测试集上达到97.1%准确率,在参考语料集上为91.2%,在ASR语料集上为90.2%。结果表明模型在真实噪声场景下具有强泛化能力。本工作建立了孟加拉语标点恢复的强基线,并公开了数据集与代码,推动低资源自然语言处理研究。

原文摘要 · Abstract (English)

Punctuation restoration enhances the readability of text and is critical for post-processing tasks in Automatic Speech Recognition (ASR), especially for low-resource languages like Bangla. In this study, we explore the application of transformer-based models, specifically XLM-RoBERTa-large, to automatically restore punctuation in unpunctuated Bangla text. We focus on predicting four punctuation marks: period, comma, question mark, and exclamation mark across diverse text domains. To address the scarcity of annotated resources, we constructed a large, varied training corpus and applied data augmentation techniques. Our best-performing model, trained with an augmentation factor of alpha = 0.20%, achieves an accuracy of 97.1% on the News test set, 91.2% on the Reference set, and 90.2% on the ASR set. Results show strong generalization to reference and ASR transcripts, demonstrating the model's effectiveness in real-world, noisy scenarios. This work establishes a strong baseline for Bangla punctuation restoration and contributes publicly available datasets and code to support future research in low-resource NLP.

标点恢复孟加拉语Transformer低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。