首个评估大模型孟加拉社交语用能力的基准,发现其常因文化误解导致表达不当。
BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social Interaction
- 构建涵盖称呼、亲属关系与习俗的三域评估基准
- 12个模型在零样本下普遍出现正式过度、称呼混淆等错误
- 特别在长幼与非正式场景中错误集中,反映系统性文化盲区
大语言模型虽具备多语言流利度,但流利不等于社交得体。在高语境语言中,沟通能力需敏感于社会等级、关系角色和互动规范,这些常直接编码于日常语言中。孟加拉语通过三层代词体系、基于亲属的称呼方式及嵌入式文化习俗体现了这一挑战。我们提出 BanglaSocialBench,首个通过情境化语言使用而非事实回忆来评估孟加拉语社会语用能力的基准。该基准覆盖三个领域:孟加拉称呼语、亲属关系推理与社会习俗,包含1,719个由母语者撰写并验证的文化相关实例。我们在零样本设置下评估12个主流大模型,观察到系统性文化错位现象:模型频繁采用过于正式的称呼形式,无法识别多种可接受的称呼代词,并在宗教背景下混淆亲属术语。研究发现,社会语用失败具有结构性且非随机;例如,不当称呼在向下层级(长辈→晚辈)和非正式语境中尤为集中。这揭示了当前大模型在真实孟加拉社交互动中推断与应用文化适切语言的持续局限。
原文摘要 · Abstract (English)
Large Language Models have demonstrated strong multilingual fluency, yet fluency alone does not guarantee socially appropriate language use. In high-context languages, communicative competence requires sensitivity to social hierarchy, relational roles, and interactional norms that are encoded directly in everyday language. Bangla exemplifies this challenge through its three-tiered pronominal system, kinship-based addressing, and culturally embedded social customs. We introduce BanglaSocialBench, the first benchmark designed to evaluate sociopragmatic competence in Bangla through context-dependent language use rather than factual recall. The benchmark spans three domains: Bangla Address Terms, Kinship Reasoning, and Social Customs, comprising 1,719 culturally grounded instances written and verified by native Bangla speakers. We evaluate twelve contemporary LLMs in a zero-shot setting and observe systematic patterns of cultural misalignment. Models frequently default to overly formal address forms, fail to recognize multiple socially acceptable address pronouns, and conflate kinship terminology across religious contexts. Our findings show that sociopragmatic failures are often structured and non-random; for example, inappropriate addressing choices concentrate heavily in downward-hierarchy (Elder$\rightarrow$Younger) and informal contexts. This reveals persistent limitations in how current LLMs infer and apply culturally appropriate language use in realistic Bangladeshi social interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。