用合成数据提升代码混用文本的语法纠错能力
LLM-based Code-Switched Text Generation for Grammatical Error Correction
- 基于合成数据构建首个大规模代码混用语法纠错数据集
- 新模型在真实代码混用文本上纠错准确率显著优于现有系统
- 适合语言学习者与多语教育技术开发者使用
随着全球化发展,跨语言混用(CSW)已成为多语对话的常见现象,给自然语言处理带来新挑战,尤其在语法错误修正(GEC)领域。本文研究了现有GEC系统在真实英语作为第二语言学习者产生的代码混用文本上的表现,探索了合成数据生成以缓解数据稀缺问题,并开发出能在单语和代码混用文本中进行语法纠错的模型。通过生成合成的代码混用GEC数据,构建了该任务中首个较大规模的数据集。实验表明,基于该数据训练的模型在真实测试集上性能显著优于现有系统。本工作面向英语学习者,旨在提供支持其英语语法提升的教育技术,同时尊重其自然的多语表达习惯。
原文摘要 · Abstract (English)
With the rise of globalisation, code-switching (CSW) has become a ubiquitous part of multilingual conversation, posing new challenges for natural language processing (NLP), especially in Grammatical Error Correction (GEC). This work explores the complexities of applying GEC systems to CSW texts. Our objectives include evaluating the performance of state-of-the-art GEC systems on an authentic CSW dataset from English as a Second Language (ESL) learners, exploring synthetic data generation as a solution to data scarcity, and developing a model capable of correcting grammatical errors in monolingual and CSW texts. We generated synthetic CSW GEC data, resulting in one of the first substantial datasets for this task, and showed that a model trained on this data is capable of significant improvements over existing systems. This work targets ESL learners, aiming to provide educational technologies that aid in the development of their English grammatical correctness without constraining their natural multilingualism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。