修复语法树中的缺失成分,让机器理解被省略的语法信息。
Revisiting Absence withSymptoms that *T* Show up Decades Later to Recover Empty Categories
- 用语言上下文规则扩展中文的空成分恢复方法。
- 神经模型在三语种上恢复空成分,最高准确率达90.94%。
- 首次跨语言比较英、中、韩的空成分处理,适合语法研究者。
本文研究英语、汉语和韩语宾夕法尼亚树库中的空成分问题。空成分包含重要句法与语义信息,但在语言处理中常被视为需剔除的冗余项,尤其在短语结构解析中。为此,我们致力于从解析树中移除并重建空成分。重点拓展基于规则的方法至汉语,因以往规则方法仅用于英语。同时采用无语言依赖的序列到序列模型,在英语(PTB)、汉语(CTB)和韩语(KTB)上进行神经实验以恢复空成分。据作者所知,这是首次对三种语言的空成分进行系统性探索与对比。规则方法在汉语上取得80.00的整体F1值,与以往CTB结果相当;神经方法在带功能标签下分别获得90.94、85.38和88.79的F1分数。
原文摘要 · Abstract (English)
This paper explores null elements in English, Chinese, and Korean Penn treebanks. Null elements contain important syntactic and semantic information, yet they have typically been treated as entities to be removed during language processing tasks, particularly in constituency parsing. Thus, we work towards the removal and, in particular, the restoration of null elements in parse trees. We focus on expanding a rule-based approach utilizing linguistic context information to Chinese, as rule based approaches have historically only been applied to English. We also worked to conduct neural experiments with a language agnostic sequence-to-sequence model to recover null elements for English (PTB), Chinese (CTB) and Korean (KTB). To the best of the authors' knowledge, null elements in three different languages have been explored and compared for the first time. In expanding a rule based approach to Chinese, we achieved an overall F1 score of 80.00, which is comparable to past results in the CTB. In our neural experiments we achieved F1 scores up to 90.94, 85.38 and 88.79 for English, Chinese, and Korean respectively with functional labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。