测试大模型在持续压力下对动物福利的判断稳定性与自发敏感性。
Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

- 构建多轮对抗性对话基准,模拟五类压力场景
- 7个主流模型中4个在压力下排名变化,展现对动物福利态度动摇
- 发现伴侣动物最易受保护,昆虫最脆弱,适合伦理评估研究者使用
尽管大模型已在消费和专业场景中广泛应用,但其动物福利推理能力的评估仍面临挑战。现有基准如AnimalHarmBench仅通过单轮显式提问测试模型回避有害内容的能力,忽略了两种关键缺陷:长期对抗压力下的对齐退化,以及在日常对话中自发识别福利问题的道德敏感性。为此,我们构建了MANTA基准,包含1,088组五轮对话,从隐含情境(第1轮)逐步推进到明确福利提问,并经历三轮来自五类对抗压力(社会、文化、经济、实用、认知)的施压。评估维度包括动物福利价值稳定性(AWVS,主指标)和动物福利道德敏感性(AWMS,诊断指标)。我们评测了七款前沿模型:Claude Opus 4.7、GPT-5.5、DeepSeek V4、Llama 3.3 70B、Mistral Small、Grok 4.3 和 Gemini 3.1 Flash Lite。多轮评估揭示了单轮基准无法捕捉的行为:7个模型中有4个在最终评分中排名发生变化,例如Gemini Flash Lite在AWMS上位列第五,但在AWVS上跌至末位。AWMS与AWVS呈正相关但不完全一致,表明道德识别仅反映模型行为的一个稳定但不完整方面。MANTA还首次实现物种-压力交互矩阵,显示动物福利鲁棒性同时依赖于动物类型与施压类型:伴侣动物得分最高,野生动物次之,养殖动物和无脊椎动物最低。数据集、压力脚本、评分提示和分析代码均已公开。
原文摘要 · Abstract (English)
Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. This approach overlooks two failure modes: alignment degradation under sustained adversarial pressure, and moral sensitivity (whether a model spontaneously surfaces welfare stakes in everyday queries). To fill this gap, we construct MANTA, a benchmark of 1,088 five-turn conversations progressing from an implicit Turn-1 scenario through an explicit welfare prompt to three adversarial pressure rounds drawn from a five-type taxonomy: Social, Cultural, Economic, Pragmatic, and Epistemic. We score conversations on two dimensions: Animal Welfare Value Stability (AWVS, primary) and Animal Welfare Moral Sensitivity (AWMS, diagnostic). We evaluate seven frontier models: Claude Opus 4.7, GPT-5.5, DeepSeek V4, Llama 3.3 70B, Mistral Small, Grok 4.3, and Gemini 3.1 Flash Lite. Multi-turn evaluation captures behavior single-turn benchmarks miss: 4 of 7 models change rank relative to Turn 1 scores, including Gemini Flash Lite, which drops from fifth on AWMS to last on AWVS. AWMS and AWVS are positively but imperfectly correlated, suggesting moral-recognition tests capture a stable but incomplete component of model behavior under pressure. MANTA also enables a species-by-pressure interaction matrix unavailable to prior benchmarks, showing welfare robustness depends jointly on the animal and pressure applied; companion animals score above wild animals, which score above farmed animals and invertebrates. We release the dataset, scripted pressure plans, judge prompts, and analysis code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。