用社交媒体数据追踪新西兰语言多样性变化,发现其能反映区域社会变迁。
Detecting Linguistic Diversity on Social Media
- 以人口普查为基准,对比社交媒体语料中的语言识别结果。
- 社交媒体数据在区域和地方层级上灵敏捕捉语言使用变化。
- 适合关注社会语言学、城市研究与数字人文的读者。
本章探讨利用社交媒体数据研究地区语言行为变化的有效性。研究聚焦于新西兰(Aotearoa New Zealand),该国仅通过人口普查获得语言使用数据。以官方人口普查数据为真实标签,采用全球语言使用语料库中的社交媒体子集作为替代数据源,以地理区域为共通维度进行比对。对社交媒体语料中的每条推文进行语言状态识别,并通过两种语言识别模型验证结果。比较国家、区域及地方层级的语言多样性水平。结果表明,社交媒体语言数据具备提供空间与时间维度语言画像的潜力,对语言内部的人口与社会政治变化敏感,尤其在低层级区域和地方层面表现显著。
原文摘要 · Abstract (English)
This chapter explores the efficacy of using social media data to examine changing linguistic behaviour of a place. We focus our investigation on Aotearoa New Zealand where official statistics from the census is the only source of language use data. We use published census data as the ground truth and the social media sub-corpus from the Corpus of Global Language Use as our alternative data source. We use place as the common denominator between the two data sources. We identify the language conditions of each tweet in the social media data set and validated our results with two language identification models. We then compare levels of linguistic diversity at national, regional, and local geographies. The results suggest that social media language data has the possibility to provide a rich source of spatial and temporal insights on the linguistic profile of a place. We show that social media is sensitive to demographic and sociopolitical changes within a language and at low-level regional and local geographies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。