用维基词典分析英语方言,发现其词库覆盖优于传统词典。
Wiktionary as a Crowdsourced Lexicon for English Dialects

- 通过两阶段方法分析12种英语方言的维基词典数据。
- 新西兰英语词形模式与牛津词典高度一致(相关系数0.883)。
- 适合研究方言、社会媒体语言或开源词典的学者使用。
本文评估维基词典作为英语方言的伦理化众包词典的可行性。采用两阶段方法,先对12种国家英语变体进行深入描述性分析,再将该词典应用于地理标记的国家级社交媒体语料,检验其在真实场景中的表现。结果表明,维基词典在区域性和外圈英语变体上的词汇覆盖率可媲美甚至超过《牛津英语词典》(OED)。针对新西兰英语的专项案例研究显示,维基词典与OED在词形构成模式上高度一致(相关系数R = 0.883)。同时,词典与地理标记社交媒体语言也表现出高一致性。尽管维基词典在词汇覆盖方面表现广泛,但研究也揭示了评估方言敏感型语言资源时面临的一些宏观挑战,如语言接触的影响及网络语料中的语域效应。
原文摘要 · Abstract (English)
This paper evaluates Wiktionary as an ethically crowdsourced lexicon for English dialects. We took a two-phase approach, providing an in-depth descriptive analysis of the crowdsourced lexicon for 12 national varieties of English before applying the lexicon to geo-referenced, country-level social media language data to examine the real-world performance of this crowdsourced dialect lexicon. We demonstrate that Wiktionary matches or exceeds the coverage of traditional dictionaries, such as the Oxford English Dictionary (OED), for regional and Outer-Circle varieties. Our dialect-specific case study on New Zealand English found high alignment between Wiktionary and the OED based on word-formation patterns (R = 0.883). Similarly, we observed high alignment between the dialect lexicon and geo-referenced social media language. While this paper found that Wiktionary has broad coverage of lexical properties, it also highlighted some of the macro-challenges involved in evaluating dialect-responsive language resources and tools, such as the role of language contact in dialects and register effects in web-based corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。