用合成变体数据提升罗马化文本的语言识别准确率
Improving Informally Romanized Language Identification
- 通过模拟自然拼写差异生成训练数据
- 在20种印地语系语言上达到88.2%的F1分数
- 适合处理拼写不规范的非拉丁文罗马化文本
拉丁字母常被非正式用于书写非拉丁文字母的语言。许多情况下(如印度多数语言),拉丁字母拼写缺乏标准,导致拼写高度不一致,使原本易区分的语言(如印地语和乌尔都语)变得高度混淆。本文通过改进训练数据的合成方法,提升罗马化文本的语言识别(LID)准确率。我们发现,使用包含自然拼写变异的合成样本训练,比使用真实存在的自然文本或训练更大模型更能提高系统性能。在Bhasha-Abhijnaanam评估集(Madhani et al., 2023a)上的20种印地语系语言测试中,性能从报告的74.7%(基于预训练神经模型)提升至85.4%(仅用合成数据训练线性分类器),进一步提升至88.2%(同时使用合成与真实采集文本)。
原文摘要 · Abstract (English)
The Latin script is often used to informally write languages with non-Latin native scripts. In many cases (e.g., most languages in India), the lack of conventional spelling in the Latin script results in high spelling variability. Such romanization renders languages that are normally easily distinguished due to being written in different scripts - Hindi and Urdu, for example - highly confusable. In this work, we increase language identification (LID) accuracy for romanized text by improving the methods used to synthesize training sets. We find that training on synthetic samples which incorporate natural spelling variation yields higher LID system accuracy than including available naturally occurring examples in the training set, or even training higher capacity models. We demonstrate new state-of-the-art LID performance on romanized text from 20 Indic languages in the Bhasha-Abhijnaanam evaluation set (Madhani et al., 2023a), improving test F1 from the reported 74.7% (using a pretrained neural model) to 85.4% using a linear classifier trained solely on synthetic data and 88.2% when also training on available harvested text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。