检测大模型对动物的物种歧视,发现其默许剥削行为但不承认其错误。
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
- 构建1003项基准测试,评估模型对物种歧视言论的识别与道德判断。
- 模型识别歧视但很少谴责,常将动物牺牲视为合理,尤其在养殖动物上。
- 模型更看重认知能力而非物种,能力相当时不偏袒人类,适合伦理研究者参考。
随着大型语言模型(LLMs)广泛应用,其伦理倾向亟需审视。基于人工智能公平性与歧视研究,本文系统考察了LLMs是否存在物种主义偏见——即基于物种归属的歧视,并探究其对非人类动物的价值判断。研究涵盖三个范式:(1) SpeciesismBench,一个包含1,003个项目的基准,用于评估对物种主义陈述的识别与道德评价;(2) 已有心理学量表,比较模型与人类参与者反应;(3) 文本生成任务,探测模型对物种主义辩解的解释或抵制。结果显示,模型能可靠识别物种主义言论,但极少谴责,常将其视为道德可接受。心理测量中,模型显性物种主义略低于人类,但在直接权衡中更倾向于救一人而非多只动物。初步推测,模型可能依据认知能力而非物种本身:当能力相当,无物种偏好;若动物被描述为更具能力,则优先于低能力人类。开放文本生成任务中,模型频繁为养殖动物的伤害辩护,却拒绝为非养殖动物辩护。这表明,尽管模型反映人类主流与进步观点的混合,仍复制了根深蒂固的动物剥削文化规范。我们主张,将非人类道德主体纳入AI公平性与对齐框架至关重要,以减少此类偏见,防止物种主义态度在人工智能系统及社会中固化。
原文摘要 · Abstract (English)
As large language models (LLMs) become more widely deployed, it is crucial to examine their ethical tendencies. Building on research on fairness and discrimination in AI, we investigate whether LLMs exhibit speciesist bias -- discrimination based on species membership -- and how they value non-human animals. We systematically examine this issue across three paradigms: (1) SpeciesismBench, a 1,003-item benchmark assessing recognition and moral evaluation of speciesist statements; (2) established psychological measures comparing model responses with those of human participants; (3) text-generation tasks probing elaboration on, or resistance to, speciesist rationalizations. In our benchmark, LLMs reliably detected speciesist statements but rarely condemned them, often treating speciesist attitudes as morally acceptable. On psychological measures, results were mixed: LLMs expressed slightly lower explicit speciesism than people, yet in direct trade-offs they more often chose to save one human over multiple animals. A tentative interpretation is that LLMs may weight cognitive capacity rather than species per se: when capacities were equal, they showed no species preference, and when an animal was described as more capable, they tended to prioritize it over a less capable human. In open-ended text generation tasks, LLMs frequently normalized or rationalized harm toward farmed animals while refusing to do so for non-farmed animals. These findings suggest that while LLMs reflect a mixture of progressive and mainstream human views, they nonetheless reproduce entrenched cultural norms around animal exploitation. We argue that expanding AI fairness and alignment frameworks to explicitly include non-human moral patients is essential for reducing these biases and preventing the entrenchment of speciesist attitudes in AI systems and the societies they influence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。