用词嵌入分析中国70年官方媒体话语,揭示社会群体认知演变。
Word Embeddings Track Social Group Changes Across 70 Years in China
- 基于时序词嵌入,多时间尺度分析1950-2019年中国官方媒体话语
- 性别与阶级表述随历史变革剧烈变化,而族裔、年龄、体型刻板印象稳定
- 首次揭示非西方语境下语言如何编码社会结构,适合关注文化变迁的研究者
语言通过词汇模式编码社会对群体的认知。尽管计算方法如词嵌入可量化分析这些模式,但现有研究主要聚焦西方语境。本文首次对1950至2019年中国国家控制媒体进行大规模计算分析,考察革命性社会转型在官方话语中如何体现。采用多时间分辨率的时序词嵌入,发现中国话语表现与西方显著不同,尤其体现在经济地位、民族和性别方面。这些表征具有不同演化动态:族裔、年龄和体貌特征的刻板印象在政治剧变中保持高度稳定,而性别与阶级表述则随历史进程发生剧烈转变。本研究推进了我们对官方话语如何通过语言编码社会结构的理解,并强调了非西方视角在计算社会科学中的重要性。
原文摘要 · Abstract (English)
Language encodes societal beliefs about social groups through word patterns. While computational methods like word embeddings enable quantitative analysis of these patterns, studies have primarily examined gradual shifts in Western contexts. We present the first large-scale computational analysis of Chinese state-controlled media (1950-2019) to examine how revolutionary social transformations are reflected in official linguistic representations of social groups. Using diachronic word embeddings at multiple temporal resolutions, we find that Chinese representations differ significantly from Western counterparts, particularly regarding economic status, ethnicity, and gender. These representations show distinct evolutionary dynamics: while stereotypes of ethnicity, age, and body type remain remarkably stable across political upheavals, representations of gender and economic classes undergo dramatic shifts tracking historical transformations. This work advances our understanding of how officially sanctioned discourse encodes social structure through language while highlighting the importance of non-Western perspectives in computational social science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。