针对韩语短文本分类,构建了融合语法结构的图模型与语义对比学习。
Linguistically Informed Graph Model and Semantic Contrastive Learning for Korean Short Text Classification
- 设计分层异构图模型,整合词素、词性、命名实体层级信息
- 在4个韩语数据集上优于基线模型,显著提升分类准确率
- 适合处理形态复杂、句序灵活的黏着语短文本任务
短文本分类(STC)因上下文信息匮乏和标注数据稀缺而具挑战性。现有方法多聚焦英语,因多数基准数据集为英文,导致现有模型很少考虑韩语的语言与结构特征,如黏着形态和灵活词序。为此,本文提出LIGRAM,一种面向韩语短文本分类的分层异构图模型。该模型在词素、词性、命名实体层面构建子图,并分层融合,以弥补短文本中上下文信息不足的问题,同时精准捕捉韩语固有的语法与语义依赖关系。此外,引入语义感知对比学习(SemCon),反映文档间的语义相似性,使模型在类别边界模糊的短文本中仍能建立更清晰的决策边界。在四个韩语短文本数据集上的实验表明,LIGRAM持续优于现有基线模型。结果验证了将语言特异性图表示与SemCon结合,是解决黏着语如韩语短文本分类的有效方案。
原文摘要 · Abstract (English)
Short text classification (STC) remains a challenging task due to the scarcity of contextual information and labeled data. However, existing approaches have pre-dominantly focused on English because most benchmark datasets for the STC are primarily available in English. Consequently, existing methods seldom incorporate the linguistic and structural characteristics of Korean, such as its agglutinative morphology and flexible word order. To address these limitations, we propose LIGRAM, a hierarchical heterogeneous graph model for Korean short-text classification. The proposed model constructs sub-graphs at the morpheme, part-of-speech, and named-entity levels and hierarchically integrates them to compensate for the limited contextual information in short texts while precisely capturing the grammatical and semantic dependencies inherent in Korean. In addition, we apply Semantics-aware Contrastive Learning (SemCon) to reflect semantic similarity across documents, enabling the model to establish clearer decision boundaries even in short texts where class distinctions are often ambiguous. We evaluate LIGRAM on four Korean short-text datasets, where it consistently outperforms existing baseline models. These outcomes validate that integrating language-specific graph representations with SemCon provides an effective solution for short text classification in agglutinative languages such as Korean.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。