用决策树优化模型,精准区分英语中that作连接词和关系代词的用法。
Refining Syntactic Distinctions Using Decision Trees: A Paper on Postnominal 'That' in Complement vs. Relative Clauses
- 基于TreeTagger重构语料库,通过决策树提升语法判别能力。
- 重新训练模型后,对that的两类用法识别准确率显著提升。
- 适合语言学研究者及语法标注工具开发者参考使用。
本研究首先测试了Helmut Schmid开发的TreeTagger英文模型在我们可用测试文件上的表现,利用该模型分析英语中的关系从句与名词补语从句。通过区分' that'作为关系代词和作为连接词的两种用法,我们采用算法对原本使用Universal Dependency框架和EWT Treebank标注的语料库进行了重新标注。随后,我们提出改进模型,通过重新训练TreeTagger,并将其与Schmid的基准模型进行对比。该过程使模型能更准确地捕捉' that'作为连接词与名词性成分的细微语法差异。同时,我们考察了训练数据集大小对TreeTagger准确率的影响,并评估了EWT Treebank文件在相关结构上的代表性。此外,还分析了影响学习效果的语言与结构因素。
原文摘要 · Abstract (English)
In this study, we first tested the performance of the TreeTagger English model developed by Helmut Schmid with test files at our disposal, using this model to analyze relative clauses and noun complement clauses in English. We distinguished between the two uses of "that," both as a relative pronoun and as a complementizer. To achieve this, we employed an algorithm to reannotate a corpus that had originally been parsed using the Universal Dependency framework with the EWT Treebank. In the next phase, we proposed an improved model by retraining TreeTagger and compared the newly trained model with Schmid's baseline model. This process allowed us to fine-tune the model's performance to more accurately capture the subtle distinctions in the use of "that" as a complementizer and as a nominal. We also examined the impact of varying the training dataset size on TreeTagger's accuracy and assessed the representativeness of the EWT Treebank files for the structures under investigation. Additionally, we analyzed some of the linguistic and structural factors influencing the ability to effectively learn this distinction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。