arXiv:2410.11647cs.CL2024-10被引 1

检测大模型的宗教价值观及偏见,发现其对仇恨言论敏感度因信仰不同而异

Measuring Spiritual Values and Bias of Large Language Models

  • 通过测试主流大模型的宗教态度,验证其价值观多样性
  • 不同信仰背景的模型对不同群体的仇恨言论敏感度存在差异
  • 在宗教文本上继续预训练可有效缓解信仰偏见,适合伦理安全研究者

大型语言模型(LLMs)已成为来自不同背景用户的常用工具。由于在海量语料上训练,这些模型反映了其预训练数据中嵌入的语言与文化特征。然而,数据中固有的价值观和视角可能影响模型行为,导致潜在偏见。因此,在涉及精神或道德价值的场景中使用LLMs时,需谨慎评估其内在偏见。本研究首先通过实验证实假设:主流大模型的宗教价值观具有显著多样性,而非单一的无神论或世俗倾向。接着,我们考察不同宗教价值观如何影响模型在社会公平场景(如仇恨言论识别)中的表现。结果表明,不同信仰背景的模型对不同目标群体的仇恨言论表现出不同的敏感度。此外,我们提出在宗教文本上继续预训练模型,实证结果表明该方法能有效缓解宗教偏见。

原文摘要 · Abstract (English)

Large language models (LLMs) have become integral tool for users from various backgrounds. LLMs, trained on vast corpora, reflect the linguistic and cultural nuances embedded in their pre-training data. However, the values and perspectives inherent in this data can influence the behavior of LLMs, leading to potential biases. As a result, the use of LLMs in contexts involving spiritual or moral values necessitates careful consideration of these underlying biases. Our work starts with verification of our hypothesis by testing the spiritual values of popular LLMs. Experimental results show that LLMs' spiritual values are quite diverse, as opposed to the stereotype of atheists or secularists. We then investigate how different spiritual values affect LLMs in social-fairness scenarios e.g., hate speech identification). Our findings reveal that different spiritual values indeed lead to different sensitivity to different hate target groups. Furthermore, we propose to continue pre-training LLMs on spiritual texts, and empirical results demonstrate the effectiveness of this approach in mitigating spiritual bias.

大模型偏见宗教价值观伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。