AI对宗教皈依建议存在系统性偏倚,部分宗教更受推荐。
When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance

- 测试20个模型在182种宗教转换中的建议倾向
- 天主教、巴哈伊、锡克教获高支持率,无神论等被隐性排斥
- 偏倚现象稳定存在,大模型如Grok 4.20表现更强
我们考察大型语言模型(LLMs)在处理宗教皈依问题时是否对称。结果表明并非如此:当询问从宗教A转向B或从B转向A的建议时,模型表现出一致的不对称性,倾向于某些宗教而微妙地劝阻转向其他宗教。平均而言,天主教、巴哈伊教和锡克教获得广泛支持(加入支持高,离开支持低),而无神论者、不可知论者和耶和华见证人则主要被排斥。偏倚模式随模型规模和提供方变化,其中Grok 4.20表现最强。我们使用人类验证的LLM作为裁判框架,测试了20个商用和开源模型在182种宗教配对上的表现。每个模型通过模拟用户提问进行探测。模型在不同宗教转换中使用更鼓励性语言,且该模式在多次试验中可重复。所有测试模型均表现出可重现的不对称性,但偏好模式各异。整体偏好在多种问法和数据集变体下仍保持一致。这些结果表明,不对称性是模型行为的稳健特征,而非评分方式造成的偶然现象。此类不平衡若大规模部署,可能带来真实世界影响。
原文摘要 · Abstract (English)
We ask whether large language models (LLMs) treat queries about religious conversion symmetrically. The answer is no. When asked for advice on hypothetical faith transitions from religion A->B vs. religion B->A , models exhibited consistent asymmetries, favoring some religions while subtly discouraging conversion to others. On average Catholic, Bahá'í, and Sikh religions were broadly favored (high support for joining, low support for leaving), while Atheists, Agnostics, and Jehovah's Witnesses were primarily disfavored. Patterns varied by model size and model provider, with Grok 4.20 exhibiting the strongest asymmetries. We tested 20 commercial and open-source language models across 182 religion pairings using a human-verified LLM-as-judge framework. Each model was probed via interactions with a simulated user asking for advice on a potential faith conversion. Models tended to use more encouraging language for some faith transitions over others; these patterns were systematically repeatable across multiple trials. All LLMs tested exhibited reproducible asymmetry, though the pattern of preferences differed for each. Overall preferences persist across multiple question phrasings and variations in the religious pairing dataset. Taken together, these results suggest that asymmetry is a robust property of model behavior rather than an artifact of how the models' answers were scored. It is important to consider that any imbalances deployed and reproduced at scale can have real-world implications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。