arXiv:2501.08441cs.CL2025-01被引 17

分析语言与文生图模型中的宗教偏见,提出检测与缓解策略。

Religious Bias Landscape in Language and Text-to-Image Models: Analysis, Detection, and Debiasing Strategies

  • 构建400个自然提示,测试多类任务中的宗教偏见。
  • 发现模型在文本与图像生成中普遍存在宗教刻板印象。
  • 验证修正提示可有效降低部分偏见,适合安全与伦理研究者。

本文系统分析了语言模型与文生图模型中的宗教偏见,涵盖开源与闭源系统。通过构建约400个自然发生的独特提示,评估模型在掩码填空、提示补全及图像生成等任务中的表现。实验揭示了特定宗教相关联的显著刻板印象与不平等倾向,且宗教偏见常与性别、年龄、国籍等人口因素交叉。研究进一步测试了针对性去偏技术,采用修正提示以减轻偏差。结果表明,当前模型在文本与图像生成中仍存在显著宗教偏见,凸显开发更具公平性的语言模型对实现全球接受度的紧迫性。

原文摘要 · Abstract (English)

Note: This paper includes examples of potentially offensive content related to religious bias, presented solely for academic purposes. The widespread adoption of language models highlights the need for critical examinations of their inherent biases, particularly concerning religion. This study systematically investigates religious bias in both language models and text-to-image generation models, analyzing both open-source and closed-source systems. We construct approximately 400 unique, naturally occurring prompts to probe language models for religious bias across diverse tasks, including mask filling, prompt completion, and image generation. Our experiments reveal concerning instances of underlying stereotypes and biases associated disproportionately with certain religions. Additionally, we explore cross-domain biases, examining how religious bias intersects with demographic factors such as gender, age, and nationality. This study further evaluates the effectiveness of targeted debiasing techniques by employing corrective prompts designed to mitigate the identified biases. Our findings demonstrate that language models continue to exhibit significant biases in both text and image generation tasks, emphasizing the urgent need to develop fairer language models to achieve global acceptability.

宗教偏见语言模型文生图去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。