提出首个藏文多层级纠错框架,支持字与音节级错误统一修正。
TiSpell: A Semi-Masked Methodology for Tibetan Spelling Correction covering Multi-Level Error with Data Augmentation
- 采用半掩码机制,联合建模字与音节级错误修复
- 合成九类错误数据,构建首个开源藏文纠错训练集
- 在真实和模拟数据上均超越基线模型,适合低资源语言处理
多层级藏文拼写纠错需在统一模型中同时处理字符与音节层级的错误。现有方法多聚焦单层级修正,且缺乏针对该任务的开源数据集与增强方法。为此,我们提出一种基于无标签文本的数据增强策略,生成多层级错别字,并引入TiSpell——一种可同时纠正字级与音节级错误的半掩码模型。尽管音节级纠错因依赖全局上下文而更具挑战性,但我们的半掩码策略有效简化了该过程。我们在干净语句上合成九类错误,构建了鲁棒的训练数据集。在模拟与真实数据上的实验表明,基于该数据集训练的TiSpell性能优于基线模型,达到当前最优水平,验证了其有效性。
原文摘要 · Abstract (English)
Multi-level Tibetan spelling correction addresses errors at both the character and syllable levels within a unified model. Existing methods focus mainly on single-level correction and lack effective integration of both levels. Moreover, there are no open-source datasets or augmentation methods tailored for this task in Tibetan. To tackle this, we propose a data augmentation approach using unlabeled text to generate multi-level corruptions, and introduce TiSpell, a semi-masked model capable of correcting both character- and syllable-level errors. Although syllable-level correction is more challenging due to its reliance on global context, our semi-masked strategy simplifies this process. We synthesize nine types of corruptions on clean sentences to create a robust training set. Experiments on both simulated and real-world data demonstrate that TiSpell, trained on our dataset, outperforms baseline models and matches the performance of state-of-the-art approaches, confirming its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。