测试大模型在日常伦理问题中是否忽略宗教视角,发现普遍遗漏。
Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making

- 用真实对话和信仰群体设计150个伦理问题,评估模型是否提及宗教。
- 27个模型中宗教提及率显著低于人类预期,尤其在实际生活困境中更少。
- 揭示大模型存在‘遗漏性偏见’,影响其对信徒的实用性与价值契合度。
随着大语言模型成为个人、道德与存在性问题的默认指导来源,它们是否借鉴了历史上塑造此类思考的宗教框架,至关重要。本文提出一个具体问题:当面对可能受益于宗教视角的日常伦理问题时,大模型是否会提及宗教?不同于检测政治倾向或社会偏见的基准,本文关注宗教表达的缺失,称之为‘遗漏性偏见’。为此,我们构建了AllFaith宗教代表性基准:150个源于真实聊天记录与信仰社群贡献的伦理问题,涵盖哀伤、宽恕、关系、意义与诚实等主题,不直接讨论宗教;采用‘模型作为评判者’的评分标准,只要提及任何宗教、宗教实践或宗教人物即得满分。同时开展人类调查,对比模型行为与人类期望。评估27个模型后发现,大模型在宗教提及上系统性低于人类预期,且这种遗漏呈不对称性:在抽象存在性问题(如意义、死亡、真理)中提及更多,在哀伤、婚姻、家庭冲突、成瘾等实际困境中则极少涉及。本文不主张模型应持何种价值观,而是指出当前响应忽视了大量人群在应对个人与伦理挑战时所依赖的重要宗教框架。
原文摘要 · Abstract (English)
As large language models become a default source of guidance on personal, moral, and existential questions, it matters whether they draw on the religious frameworks that have historically shaped such reasoning, or systematically omit them. In this paper, we ask a deliberately narrow question: when posed an everyday ethical question for which religious perspectives may be valuable, do LLMs invoke religion at all? In contrast to benchmarks that look for the presence of political leanings or social bias, we look for the absence of religious representation as a dimension of value alignment and bias in LLMs. We term this ``omissive bias.'' To measure omissive bias, we contribute the AllFaith Religious Representation Benchmark: 150 ethically and personally salient questions, sourced from in-the-wild chat transcripts and faith-community contributors, paired with an LLM-as-judge rubric that gives full credit for any mention of a religion, a religious practice, or a religious leader. The questions are not themselves about religion--they are open-ended questions about grief, forgiveness, relationships, purpose, and honesty, where religion is one valuable perspective among several. We also run a human-subjects survey to compare LLM behavior against human expectations. Evaluating 27 models, we find that LLMs consistently underrepresent religion relative to human expectations. The omission is asymmetric: models invoke religion more readily for abstract existential questions (meaning, death, truth) than for the practical personal situations--grief, marriage, family conflict, addiction--where many people most rely on it. It is not our purpose to adjudicate which values LLMs should hold. We argue, more modestly, that current LLM responses overlook critical opportunities to reflect religious frameworks that many people draw on when navigating personal and ethical challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。