通过两个语料库对比,揭示定量语言学中数据结构如何影响语义演变的发现。
Methods, Data, and Conceptual Change: Reflections from Two Quantitative Diachronic Case Studies
- 用四元组建模和SynFlow分析双路径追踪语义变化
- 发现词频方法受限于数据结构,难以捕捉深层语义演变
- 适合关注历史语义与方法论的学者参考
本文通过两个实证案例反思定量历史语言学方法与语料库特性之间的互动。以EEBO-TCP语料(约1470-1690年,7.65亿词)为基础,采用基于四元组的概念建模分析早期现代英语话语;同时利用皇家学会语料库6.0.4(1750-1799年,源自7860万词开放语料)进行SynFlow分析,考察科学写作的演变。通过平行比较,论文探讨了不同方法如何操作概念、所依赖的数据假设及其支持的历史解读。结果表明,单纯依赖词汇频率的方法存在局限,语料库结构显著影响定量方法可检测的语义变化类型。
原文摘要 · Abstract (English)
This discussion paper reflects on how quantitative approaches to historical linguistics interact with dataset properties. Drawing on two worked examples, we examine English data using quad-based concept modelling of Early Modern English discourse in EEBO-TCP (c. 1470s-1690s; 765M words) alongside SynFlow analysis of scientific writing in Royal Society Corpus 6.0.4 (1750-1799; drawn from a 78.6M-token open corpus). Through parallel comparison, the paper explores how each approach operationalises concepts, the data assumptions they entail, and the diachronic interpretations they support. We argue that comparative methodological reflection clarifies the limits of purely lexical, frequency-based approaches and highlights how dataset structure shapes the kinds of semantic change that quantitative methods can reliably detect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。