arXiv:2504.04640cs.CLcs.AI2025-04ACL

构建可扩展的社交文化语言研究平台,支持快速验证语言差异假设。

Splits! Flexible Sociocultural Linguistic Investigation at Scale

  • 通过拆分Reddit数据集实现人口与话题维度的灵活分析
  • 验证了多个已有文化语言模式,如中美饮食观念差异
  • 提供两阶段流程筛选潜在语言现象,适合快速探索新假设

语言使用中的差异受说话者社会文化背景和具体语境影响,是理解文化视角、价值观与观点的丰富窗口。例如,中国学生谈论“健康饮食”时多用“时间”“规律”“消化”等词,而美国人则倾向“平衡食物组”“避免脂肪糖分”,反映出不同的营养认知模型。传统上,计算层面的社会文化语言现象(SLP)研究依赖特定群体或主题的定制化分析,需专门的数据收集与实验设计,难以支持快速假设检验与原型开发。为此,我们提出构建一个面向系统性、灵活性社会语言学研究的“沙盒”框架。利用该方法,我们构建了基于人口统计与话题划分的Reddit数据集Splits!,并通过自我认同信息与复现现有文献中的多个已知SLP进行验证。我们进一步展示该沙盒在规模化两阶段流程中的应用:从大量潜在社会文化语言现象(PSLPs)中筛选出最值得深入质性分析的候选项。

原文摘要 · Abstract (English)

Variation in language use, shaped by speakers' sociocultural background and specific context of use, offers a rich lens into cultural perspectives, values, and opinions. For example, Chinese students discuss "healthy eating" with words like "timing," "regularity," and "digestion," whereas Americans use vocabulary like "balancing food groups" and "avoiding fat and sugar," reflecting distinct cultural models of nutrition. The computational study of these Sociocultural Linguistic Phenomena (SLP) has traditionally been done in NLP via tailored analyses of specific groups or topics, requiring specialized data collection and experimental operationalization--a process not well-suited to quick hypothesis exploration and prototyping. To address this, we propose constructing a "sandbox" designed for systematic and flexible sociolinguistic research. Using our method, we construct a demographically/topically split Reddit dataset, Splits!, validated by self-identification and by replicating several known SLPs from existing literature. We showcase the sandbox's utility with a scalable, two-stage process that filters large collections of "potential" SLPs (PSLPs) to surface the most promising candidates for deeper, qualitative investigation.

社会语言学文化差异数据分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。