arXiv:2602.11939cs.CL2026-02

大模型对不同社会阶层语言风格适应能力弱,易强化语言不平等

Do Large Language Models Adapt to Language Variation across Socioeconomic Status?

  • 用社交媒体数据测试大模型对不同阶层语言的模仿能力
  • 模型对高社会阶层语言模仿更好,低阶层语言常被歪曲或简化
  • 提醒研究者慎用大模型做社会模拟或语言风格分析

人类会根据交流对象调整语言风格,但大语言模型(LLM)在多大程度上能适应不同社会背景尚不清楚。随着这些模型越来越多地介入人与人之间的沟通,若无法适配多元语言风格,可能加剧偏见并边缘化语言规范与模型不匹配的群体,从而固化社会分层。本文基于Reddit和YouTube的全新数据集,按社会经济地位(SES)分层,使用四个LLM对语料库中不完整文本进行补全,并与原始文本在94项社会语言学指标(包括句法、修辞、词汇特征)上对比。结果显示,LLM对不同SES的语言风格调节能力极有限,常导致近似或刻板化表达,且更擅长模仿高SES语言。研究揭示了大模型可能放大语言等级制度的风险,并质疑其在基于代理的社会模拟、调查实验及依赖语言风格作为社会信号的研究中的适用性。

原文摘要 · Abstract (English)

Humans adjust their linguistic style to the audience they are addressing. However, the extent to which LLMs adapt to different social contexts is largely unknown. As these models increasingly mediate human-to-human communication, their failure to adapt to diverse styles can perpetuate stereotypes and marginalize communities whose linguistic norms are less closely mirrored by the models, thereby reinforcing social stratification. We study the extent to which LLMs integrate into social media communication across different socioeconomic status (SES) communities. We collect a novel dataset from Reddit and YouTube, stratified by SES. We prompt four LLMs with incomplete text from that corpus and compare the LLM-generated completions to the originals along 94 sociolinguistic metrics, including syntactic, rhetorical, and lexical features. LLMs modulate their style with respect to SES to only a minor extent, often resulting in approximation or caricature, and tend to emulate the style of upper SES more effectively. Our findings (1) show how LLMs risk amplifying linguistic hierarchies and (2) call into question their validity for agent-based social simulation, survey experiments, and any research relying on language style as a social signal.

大模型语言风格社会偏见社会阶层

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。