扩充韩语学习者语料库,支持更精准的写作评估与自动纠错。
Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring
- 为语料库添加多参考纠错标注,反映语言多样性。
- 引入韩国国家语言院评分标准,涵盖语法、连贯性与词汇多样性。
- 适合韩语教学、自动纠错与语言测评研究者使用。
尽管全球对韩语教育的兴趣日益增长,但针对韩语第二语言写作的学习者语料库仍严重不足。为弥补这一缺口,我们通过添加多个语法错误纠正(GEC)参考答案,增强了KoLLA韩语学习者语料库,从而支持对GEC系统更细致、灵活的评估,并体现人类语言的变异性。此外,我们依据韩国国家语言院指南,为语料库增添了基于评分量表的评分,涵盖语法准确性、连贯性和词汇多样性。这些改进使KoLLA成为韩语二语教育研究中可靠且标准化的资源,推动语言学习、评估及自动纠错技术的发展。
原文摘要 · Abstract (English)
Despite growing global interest in Korean language education, there remains a significant lack of learner corpora tailored to Korean L2 writing. To address this gap, we enhance the KoLLA Korean learner corpus by adding multiple grammatical error correction (GEC) references, thereby enabling more nuanced and flexible evaluation of GEC systems, and reflects the variability of human language. Additionally, we enrich the corpus with rubric-based scores aligned with guidelines from the Korean National Language Institute, capturing grammatical accuracy, coherence, and lexical diversity. These enhancements make KoLLA a robust and standardized resource for research in Korean L2 education, supporting advancements in language learning, assessment, and automated error correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。