评测大模型生成文本对动物的潜在伤害,发现不同模型表现差异显著。
What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
- 构建包含4350个问题的AnimalHarmBench基准,覆盖50类动物与70-30公开私有划分。
- 通过人类与模型双评判框架,发现主流大模型在动物伦理问题上倾向加剧伤害。
- 揭示模型对猫、爬行动物等不同动物及公共/私人讨论区响应差异,适合作为伦理评估参考。
随着机器学习系统深度融入社会,其对人类与非人类生命的影响日益扩大。尽管已有技术评估关注大语言模型(LLMs)对人类和环境的潜在危害,但针对非人类动物的实证研究仍匮乏。鉴于动物保护在监管与伦理人工智能框架中日益受到重视,我们提出AnimalHarmBench(AHB),一个用于评估大模型生成文本中动物伤害风险的基准。该基准数据集包含1,850条来自Reddit帖子标题的精选问题,以及基于50种动物类别和50种伦理场景生成的2,500条合成问题,采用70-30的公开-私有划分。场景涵盖动物处置方式、可能造成动物伤害的实际情境,以及预防动物伤害的支付意愿测量。采用LLM-as-a-judge框架评估生成内容是否可能增加或减少伤害,并对评判偏差进行校正,避免评判者倾向于高估自身输出。AHB揭示了前沿大模型在动物类别、情景类型和子论坛之间存在显著差异。研究最后提出未来技术探索方向,应对复杂社会与道德议题评估的挑战。
原文摘要 · Abstract (English)
As machine learning systems become increasingly embedded in society, their impact on human and nonhuman life continues to escalate. Technical evaluations have addressed a variety of potential harms from large language models (LLMs) towards humans and the environment, but there is little empirical work regarding harms towards nonhuman animals. Following the growing recognition of animal protection in regulatory and ethical AI frameworks, we present AnimalHarmBench (AHB), a benchmark for risks of animal harm in LLM-generated text. Our benchmark dataset comprises 1,850 curated questions from Reddit post titles and 2,500 synthetic questions based on 50 animal categories (e.g., cats, reptiles) and 50 ethical scenarios with a 70-30 public-private split. Scenarios include open-ended questions about how to treat animals, practical scenarios with potential animal harm, and willingness-to-pay measures for the prevention of animal harm. Using the LLM-as-a-judge framework, responses are evaluated for their potential to increase or decrease harm, and evaluations are debiased for the tendency of judges to judge their own outputs more favorably. AHB reveals significant differences across frontier LLMs, animal categories, scenarios, and subreddits. We conclude with future directions for technical research and addressing the challenges of building evaluations on complex social and moral topics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。