通过树库对比口语与书面语的句法差异,发现两者结构类型差异显著。
Counting trees: A treebank-driven exploration of syntactic variation in speech and writing across languages
- 用去词化的依存子树定义句法结构,从英斯语树库中提取分析。
- 口语结构数量少、多样性低,且与书面语重叠率极低。
- 适合研究语言使用中的句法变异,尤其关注互动与表达经济性。
本文提出一种基于树库的新型方法,比较口语与书面语中的句法结构,使用依赖标注的语料库。采用完全归纳、自下而上的方式,将句法结构定义为去词化的依存(子)树,并从英语和斯洛文尼亚语的通用依存(UD)树库中提取口语与书面语数据。对每个语料库,分析句法库存的规模、多样性和分布特征,以及不同语域间的重叠程度,识别出最具口语性的结构。结果表明:在两种语言中,口语语料的句法结构更少、多样性更低,且跨语域的结构重叠极为有限——多数口语结构在书面语中不出现,反映出实时互动与精心写作在句法组织上的特定需求。关键性分析进一步揭示,高频口语特有结构具有互动性、语境锚定和表达经济等特点。该可扩展、语言无关的框架为系统研究语料间句法变异提供了有效路径,为基于数据驱动的语法使用理论奠定基础。
原文摘要 · Abstract (English)
This paper presents a novel treebank-driven approach to comparing syntactic structures in speech and writing using dependency-parsed corpora. Adopting a fully inductive, bottom-up method, we define syntactic structures as delexicalized dependency (sub)trees and extract them from spoken and written Universal Dependencies (UD) treebanks in two syntactically distinct languages, English and Slovenian. For each corpus, we analyze the size, diversity, and distribution of syntactic inventories, their overlap across modalities, and the structures most characteristic of speech. Results show that, across both languages, spoken corpora contain fewer and less diverse syntactic structures than their written counterparts, with consistent cross-linguistic preferences for certain structural types across modalities. Strikingly, the overlap between spoken and written syntactic inventories is very limited: most structures attested in speech do not occur in writing, pointing to modality-specific preferences in syntactic organization that reflect the distinct demands of real-time interaction and elaborated writing. This contrast is further supported by a keyness analysis of the most frequent speech-specific structures, which highlights patterns associated with interactivity, context-grounding, and economy of expression. We argue that this scalable, language-independent framework offers a useful general method for systematically studying syntactic variation across corpora, laying the groundwork for more comprehensive data-driven theories of grammar in use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。