用Transformer模型补全元音音节文字的残缺输入,提升文本预测准确率。
Reconstructing Syllable Sequences in Abugida Scripts with Incomplete Inputs
- 基于Transformer建模,从残缺输入中重建完整音节序列。
- 辅音信息对预测帮助大,拼写错误修复效果显著优于元音缺失场景。
- 适用于拼写纠错、输入法预测等实际应用,尤其适合东南亚音节文字。
本文研究基于Transformer模型在六种阿布吉达语言(孟加拉语、印地语、高棉语、老挝语、缅甸语、泰语)中的音节序列预测问题,数据来自亚洲语言语料库(ALT)。针对辅音序列、元音序列、部分音节(随机字符删除)及掩码音节(固定音节删除)等不完整输入形式,进行音节序列重建。实验表明,辅音序列对准确预测起关键作用,取得较高BLEU分数;而元音序列则更具挑战性。模型在处理部分音节与掩码音节重建任务中表现稳健,尤其在包含辅音信息和音节掩码的任务中结果突出。本研究深化了对阿布吉达语言序列预测的理解,并为文本预测、拼写纠错与数据增强提供了实用参考。
原文摘要 · Abstract (English)
This paper explores syllable sequence prediction in Abugida languages using Transformer-based models, focusing on six languages: Bengali, Hindi, Khmer, Lao, Myanmar, and Thai, from the Asian Language Treebank (ALT) dataset. We investigate the reconstruction of complete syllable sequences from various incomplete input types, including consonant sequences, vowel sequences, partial syllables (with random character deletions), and masked syllables (with fixed syllable deletions). Our experiments reveal that consonant sequences play a critical role in accurate syllable prediction, achieving high BLEU scores, while vowel sequences present a significantly greater challenge. The model demonstrates robust performance across tasks, particularly in handling partial and masked syllable reconstruction, with strong results for tasks involving consonant information and syllable masking. This study advances the understanding of sequence prediction for Abugida languages and provides practical insights for applications such as text prediction, spelling correction, and data augmentation in these scripts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。