为中文学习者语法错误标注设计分层体系,兼顾自动纠错与教学分析。
A Layered Taxonomy for Chinese Learner Grammatical Error Annotation
- 分层标注:先标错字标点,再用操作+领域+词性三重标签
- 覆盖率达98.7%的自动纠错数据,五模型标注一致性良好
- 适配中文特有语法结构,适合语言教学与纠错系统开发
中文学习者写作中的语法错误标注需要兼具一致性和语言学意义。本文提出一种分层标注体系,连接计算型中文语法纠错(CGEC)与教学型错误分析。该体系首先识别字符和标点层面的拼写错误,按编辑操作与子类型标注;其他错误采用三层核心标签,包含编辑操作、语言领域与词性,并可选添加中文特有的体貌、情态、比较、论元结构及补语等扩展。基于CGEC资源、学习者错误分类体系及普通话语法,通过分析自动生成的MuCGEC编辑数据覆盖情况,以及五种大语言模型在样本上的初步标注一致性研究,验证了分层方法的有效性,同时发现部分类别边界仍需细化。
原文摘要 · Abstract (English)
Grammatical error annotation in Chinese learner writing requires labels that are both consistent and linguistically meaningful. This paper proposes a layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis. The scheme first identifies character- and punctuation-level orthographic errors, labeling them by edit operation and subtype. Other errors receive a three-layer core label combining edit operation, linguistic domain, and part of speech, with optional Chinese-specific extensions for aspect, modality, comparison, argument structure, and complements. Drawing on CGEC resources, learner-error taxonomies, and Mandarin grammar, the taxonomy is evaluated through a coverage analysis of automatically extracted MuCGEC edits and a preliminary consistency study in which five large language models apply it to a sample. The results support the layered approach while identifying category boundaries requiring further refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。