arXiv:2605.29638cs.CL2026-05

构建韩语非正式文本分类模型,提升在线学习系统真实语料处理能力

Classification of non-analyzable word types in web documents to implement an effective Korean e-learning system

  • 对比正式新闻与非正式博客文本,识别语体差异特征
  • 发现非正式文本占比较高,需针对性建模处理
  • 提出局部语法图(LGG)模型,适配韩语非正式表达

在线学习系统应呈现语言实际使用中的多样现象。除正式韩语外,融入网络文档、短信或推文等真实语境表达,对高级学习者尤为有益。本文构建两类语料库:一类为新闻文章等正式文档;另一类为网络博客中关于新产品的用户评论等非正式文档。通过对比分析,揭示两类文本在表达上的显著差异。针对非正式文本占比高的现状,提出局部语法图(Local Grammar Graphs, LGG)模型,以更有效处理韩语非正式表达,增强在线学习系统的实用性与适应性。

原文摘要 · Abstract (English)

E-learning systems should deliver contents that reflect various phenomena of the language as it is used. In addition to formal Korean, e-learning systems that would include real-world Korean expressions such as those in web documents, mobile text messages, or twitter posts, would be useful to high-level learners. We construct two types of corpora: one is made of formal documents like online news articles; the other is made of informal documents like customer reviews about new products in web blogs. By comparing these corpora, we show how expressions differ in these two types of corpora. We survey the main characteristics of the informal corpus. Given that a significant proportion of text is informal, we propose Local Grammar Graphs (LGG) as an appropriate model to treat them effectively in Korean e-learning systems.

韩语学习语料分析自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。