构建俄语母语者英语作文中的干扰错误数据集与生成框架。
RILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts
- 用规则与神经方法合成俄语母语者的英语错误,结合专家标注数据。
- 在1.8万句数据上训练模型,对拼写与时态错误识别效果显著提升。
- 适合语言教学研究者和英语学习工具开发者使用。
学生作文中的许多错误可归因于母语(L1)影响。本文聚焦俄语母语者在写作英语时受母语干扰产生的错误,如将 stadium 写成 stadion,体现词汇直译特征。我们提出 RILEC,一个包含超过 18,000 句的大型数据集,融合了 REALEC 的专家标注数据与通过规则和神经增强生成的合成样本。设计了一种基于生成式语言模型的错误生成框架,采用 PPO 优化、提示控制及规则模式实现精准干扰错误生成。在 RILEC 上微调的模型在词级干扰类型(如拼写转写、时态语义)上表现优异。实验表明,该增强流程显著提升模型性能,为学习者与教师提供高效识别与纠正此类错误的工具。
原文摘要 · Abstract (English)
Many errors in student essays can be explained by influence from the native language (L1). L1 interference refers to errors influenced by a speaker's first language, such as using stadion instead of stadium, reflecting lexical transliteration from Russian. In this work, we address the task of detecting such errors in English essays written by Russian-speaking learners. We introduce RILEC, a large-scale dataset of over 18,000 sentences, combining expert-annotated data from REALEC with synthetic examples generated through rule-based and neural augmentation. We propose a framework for generating L1-motivated errors using generative language models optimized with PPO, prompt-based control, and rule-based patterns. Models fine-tuned on RILEC achieve strong performance, particularly on word-level interference types such as transliteration and tense semantics. We find that the proposed augmentation pipeline leads to a significant performance improvement, making it a potentially valuable tool for learners and teachers to more effectively identify and address such errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。