用哈明距离量化大模型在亚洲国家的宗教偏见,发现多数模型输出同质化。
Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations
- 通过哈明距离比较模型回答与调查数据,推断其反映的人口统计特征。
- 在印度等亚洲国家测试显示,多数开放模型呈现单一、均质的宗教态度。
- 揭示模型可能传播霸权世界观,适合关注文化敏感性的研究者参考。
大型语言模型(LLMs)可能在不知情的情况下生成观点并传播偏见,根源在于数据收集缺乏代表性与多样性。以往研究主要聚焦西方,尤其是美国,但其结论未必适用于非西方群体。随着大模型在各领域广泛应用,其输出的文化敏感性至关重要。本文提出一种新方法,通过哈明距离衡量模型响应与调查受访者之间的差异,以定量分析模型生成观点所反映的社会人口特征。我们在多个全球南方国家(重点为印度及其它亚洲国家)的调查数据上评估了 Llama、Mistral 等现代开放模型,特别关注宗教包容性与身份认同议题。结果表明,大多数开放模型表现出单一且同质化的社会人口特征,且因国家/地区而异。这引发对模型可能传播主导世界观、削弱少数群体视角的风险的担忧。该框架还可用于未来研究训练数据、模型架构与输出偏见之间复杂关系,尤其在宗教包容性与身份认同等敏感话题上。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are capable of generating opinions and propagating bias unknowingly, originating from unrepresentative and non-diverse data collection. Prior research has analysed these opinions with respect to the West, particularly the United States. However, insights thus produced may not be generalized in non-Western populations. With the widespread usage of LLM systems by users across several different walks of life, the cultural sensitivity of each generated output is of crucial interest. Our work proposes a novel method that quantitatively analyzes the opinions generated by LLMs, improving on previous work with regards to extracting the social demographics of the models. Our method measures the distance from an LLM's response to survey respondents, through Hamming Distance, to infer the demographic characteristics reflected in the model's outputs. We evaluate modern, open LLMs such as Llama and Mistral on surveys conducted in various global south countries, with a focus on India and other Asian nations, specifically assessing the model's performance on surveys related to religious tolerance and identity. Our analysis reveals that most open LLMs match a single homogeneous profile, varying across different countries/territories, which in turn raises questions about the risks of LLMs promoting a hegemonic worldview, and undermining perspectives of different minorities. Our framework may also be useful for future research investigating the complex intersection between training data, model architecture, and the resulting biases reflected in LLM outputs, particularly concerning sensitive topics like religious tolerance and identity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。