arXiv:2604.12377cs.CLcs.AI2026-04ACL

为韩语模型注入字形结构知识,提升理解与生成能力

SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models

  • 通过子字符组合模块注入韩文字形结构信息
  • 在多个韩语任务上显著提升基线模型性能
  • 适合需要精细语言结构建模的韩语研究者

韩语是一种形态丰富的语言,其书写系统由称为‘音素’(Jamo)的子字符单元构成。这些子字符不仅决定韩文的视觉结构,还编码了频繁且具有语言学意义的音位形态过程。然而,当前大多数韩语语言模型基于子词分词方案,未能显式捕捉字符内部的组合结构。为此,我们提出SCRIPT——一种模型无关的模块,可将子字符组合知识注入韩语预训练语言模型中。SCRIPT无需修改架构或额外预训练,即可增强子词嵌入的结构粒度。实验表明,SCRIPT在多种韩语自然语言理解(NLU)和生成(NLG)任务中均显著提升所有基线模型表现。此外,详细的语言学分析显示,SCRIPT重塑了嵌入空间,更准确地捕捉语法规律与语义连贯性变化。代码已开源:https://github.com/SungHo3268/SCRIPT。

原文摘要 · Abstract (English)

Korean is a morphologically rich language with a featural writing system in which each character is systematically composed of subcharacter units known as Jamo. These subcharacters not only determine the visual structure of Korean but also encode frequent and linguistically meaningful morphophonological processes. However, most current Korean language models (LMs) are based on subword tokenization schemes, which are not explicitly designed to capture the internal compositional structure of characters. To address this limitation, we propose SCRIPT, a model-agnostic module that injects subcharacter compositional knowledge into Korean PLMs. SCRIPT allows to enhance subword embeddings with structural granularity, without requiring architectural changes or additional pre-training. As a result, SCRIPT enhances all baselines across various Korean natural language understanding (NLU) and generation (NLG) tasks. Moreover, beyond performance gains, detailed linguistic analyses show that SCRIPT reshapes the embedding space in a way that better captures grammatical regularities and semantically cohesive variations. Our code is available at https://github.com/SungHo3268/SCRIPT.

韩语处理子字符语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。