从大规模语料中自动提取真实英语从句,解决人工构造数据的局限性。
Automatic Extraction of Clausal Embedding Based on Large-Scale English Text Data
- 基于句法解析与解析启发式规则,自动识别自然文本中的嵌入从句。
- 在手标数据集GECS上验证效果,实现对真实语言现象的精准捕获。
- 构建了从Dolma语料库提取的大规模真实嵌入从句数据集,供后续研究使用。
对于语言学家而言,嵌入从句因其复杂的句法与语义特征分布而备受关注。然而,现有研究多依赖人为构造的语言例句来分析此类结构,忽略了从大规模语言语料中获取的统计信息和真实语例。为此,我们提出一种方法论,利用成分句法解析与一组解析启发式规则,在大规模文本数据中检测并标注自然出现的英语嵌入从句。我们的工具在自建的手工标注数据集黄金嵌入从句集(GECS)上进行了评估。最后,我们基于开源语料库Dolma,利用该提取工具构建了一个大规模的真实英语嵌入从句数据集。
原文摘要 · Abstract (English)
For linguists, embedded clauses have been of special interest because of their intricate distribution of syntactic and semantic features. Yet, current research relies on schematically created language examples to investigate these constructions, missing out on statistical information and naturally-occurring examples that can be gained from large language corpora. Thus, we present a methodological approach for detecting and annotating naturally-occurring examples of English embedded clauses in large-scale text data using constituency parsing and a set of parsing heuristics. Our tool has been evaluated on our dataset Golden Embedded Clause Set (GECS), which includes hand-annotated examples of naturally-occurring English embedded clause sentences. Finally, we present a large-scale dataset of naturally-occurring English embedded clauses which we have extracted from the open-source corpus Dolma using our extraction tool.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。