用翻译数据增强低资源语言文本难度评估,提升模型精度。
Translation as Augmentation: Effect of Translated Data on Assessment of Difficulty

- 用机器翻译将高资源语言标注数据迁移到低资源语言
- 翻译数据加入后模型预测准确率显著提升
- 适合缺乏专家标注的低资源语言研究者使用
可靠的文本难度评估是文本简化流程和个性化学习应用的前提。然而,由于缺乏细粒度难度标注(如CEFR等级)的专家语料库,尤其在低资源语言中,这一任务面临严峻挑战。本文针对一种低资源欧洲语言,提出一种跨语言数据增强策略:利用机器翻译将高资源语言的标注资源迁移至目标语言。我们训练基于BERT的回归模型预测难度分数,并探究合成翻译数据是否能有效补充原生训练集。实验表明,将稀缺原生数据与机器翻译语料结合,可显著提升难度估计的准确性,为缺乏大量专家标注的语言提供了可行解决方案。
原文摘要 · Abstract (English)
Reliable Text Difficulty Assessment is a prerequisite for valid text simplification workflows and personalized learning applications. However, the development of robust assessment models is severely hindered by a critical bottleneck: the scarcity of expert-annotated corpora containing fine-grained difficulty levels (e.g., CEFR), particularly for lower-resource languages. This paper addresses this data scarcity problem in the context of a low-resource European language. We propose a cross-lingual data augmentation strategy that leverages machine translation to transfer labeled resources from high-resource languages to the target low-resource language. We train BERT-based regression models to predict difficulty scores and investigate whether synthetic, translated data can effectively supplement native training sets. Our experiments demonstrate that augmenting scarce native data with machine-translated corpora significantly improves the accuracy of difficulty estimation, offering a viable solution for languages lacking extensive expert annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。