arXiv:2501.03479cs.CL2025-01EMNLP被引 2

对比印地语和孟加拉语维基百科与大模型对尊称的使用差异。

Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi

  • 分析1万篇维基文章中尊称用法,关联性别、年龄、名气等社会属性。
  • 孟加拉语尊称更普遍,男性比女性更常被尊称,臭名昭著或异域人物多不用尊称。
  • 大模型在尊称选择上与维基存在差异,反映其社会语言规范学习偏差。

南亚多种语言强制使用第三人称尊称,体现权力、年龄、性别、名气和社会距离等细微语用信息。本文(一)首次对10,000篇印地语和孟加拉语维基百科文章中的尊称代词与动词用法进行大规模研究,标注了主体的关键社会人口属性,包括性别、年龄组、知名度和文化起源。(二)分析发现语言内部有系统性规律,但跨语言差异显著:孟加拉语尊称使用率高于印地语;而对臭名昭著、未成年及文化异域个体,非尊称占主导。尤其在印地语中,男性被尊称频率显著高于女性。(三)为检验大语言模型(LLMs)是否内化类似社会语用规范,我们对六款大模型在1,000个文化均衡实体上执行受控生成与翻译任务。结果显示,大模型在尊称选择上偏离维基用法,不同任务、语言及社会属性下表现出替代偏好。这些差异揭示了大模型在社会文化适配上的差距,为研究大模型如何习得、适应或扭曲社会语言规范开辟新路径。代码与数据已公开于 https://github.com/souro/honorific-wiki-llm。

原文摘要 · Abstract (English)

The obligatory use of third-person honorifics is a distinctive feature of several South Asian languages, encoding nuanced socio-pragmatic cues such as power, age, gender, fame, and social distance. In this work, (i) We present the first large-scale study of third-person honorific pronoun and verb usage across 10,000 Hindi and Bengali Wikipedia articles with annotations linked to key socio-demographic attributes of the subjects, including gender, age group, fame, and cultural origin. (ii) Our analysis uncovers systematic intra-language regularities but notable cross-linguistic differences: honorifics are more prevalent in Bengali than in Hindi, while non-honorifics dominate while referring to infamous, juvenile, and culturally exotic entities. Notably, in both languages, and more prominently in Hindi, men are more frequently addressed with honorifics than women. (iii) To examine whether large language models (LLMs) internalize similar socio-pragmatic norms, we probe six LLMs using controlled generation and translation tasks over 1,000 culturally balanced entities. We find that LLMs diverge from Wikipedia usage, exhibiting alternative preferences in honorific selection across tasks, languages, and socio-demographic attributes. These discrepancies highlight gaps in the socio-cultural alignment of LLMs and open new directions for studying how LLMs acquire, adapt, or distort social-linguistic norms. Our code and data are publicly available at https://github.com/souro/honorific-wiki-llm

社会语言学大模型偏见南亚语言尊称研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。