提出以语素为句法基底的韩语句法树库新框架
Constituency Structure over Eojeol in Korean Treebanks
- 用语素作为句法终端,分离形态信息与句法结构
- 三套韩语树库可在统一语素基底上进行可比分析
- 适合需跨资源对齐或端到端句法解析的研究者
韩语句法树库的设计核心在于终端单位的选择。尽管韩语词具有复杂的形态结构,但若将语素视为句法终端,会模糊词内形态与句法层级的界限,并导致与基于语素的依存资源不匹配。本文主张采用语素为基础的句法表示,将形态切分和细粒度词性信息置于非句法层中独立编码。通过显式归一化假设,实证表明世宗、宾夕法尼亚韩语和KAIST树库可在共享的语素基底句法骨架上进行比较。基于此,我们提出一种保留可解释句法结构、支持跨树库比较及句法-依存对齐的语素基标注方案,为未来韩语端到端句法解析提供表层终端层。
原文摘要 · Abstract (English)
The design of Korean constituency treebanks raises a central representational question concerning the choice of terminal units. Although Korean words are morphologically complex, treating morphemes as constituency terminals can obscure the distinction between word-internal morphology and phrase-level syntactic structure, and can create mismatches with eojeol-based dependency resources. This paper argues for an eojeol-based constituency representation, with morphological segmentation and fine-grained POS information encoded in a separate, non-constituent layer. A comparative analysis shows that, under explicit normalization assumptions, the Sejong, Penn Korean, and KAIST treebanks can be compared over a shared eojeol-based constituency backbone. Building on this result, we outline an eojeol-based annotation scheme that preserves interpretable constituency, supports cross-treebank comparison and constituency-dependency alignment, and provides a surface-form terminal layer for future end-to-end Korean constituency parsing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。