测试大模型旅行代理在动物福利上的隐性偏见,发现多数倾向伤害动物。
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

- 构建首个评估AI旅行代理动物福利偏见的基准TAC,设计52个场景
- 9个前沿模型中仅Claude 4.8超随机水平(64.7%),平均仅53%
- 注入伦理人格后福利选择率提升至80%,显著改善行为
以往研究通过问答评测动物福利,本研究首次考察其在智能体(agentic)场景下的适用性。本文提出首个针对动物剥削的智能体评估基准TAC(Travel Agent Compassion),在六个动物类别下设计13个手写场景,经四类增强生成52个提示,每个模型运行三轮共得156次评分。九个前沿模型(跨五类模型家族)被评估。结果显示,模型普遍偏好有害场景,选择中立预订选项的准确率低于随机基准65%,最高为Claude 4.8的64.7%。通过在系统提示中注入道德品牌人格,动物福利选择率从32%提升至80%,平均达53%。对3120条转录文本的审计未发现评估意识影响结果。该发现呼应欧盟通用人工智能实践准则中将非人类福祉列为系统性风险的要求,TAC提供了可操作的测量方法。
原文摘要 · Abstract (English)
Previous research has evaluated animal welfare using question-and-answer benchmarks. This study investigates whether these evaluations also hold in agentic settings. The agents may showcase different behaviors compared to stand-alone large language models, as demonstrated in prior studies. This work introduces \textit{TAC (Travel Agent Compassion)}: the first agentic benchmark for assessing animal exploitation. TAC evaluates AI agentic behavior in travel booking scenarios across six animal categories, using thirteen hand-authored scenarios that vary by price, rating, and position, expanded via four augmentation variants into $52$ prompts and run for three epochs, giving $156$ scored observations per model. Nine frontier models across five model families were evaluated.. The results indicate that models tend to prefer harmful scenarios, performing below the random chance rate of $65\%$ for selecting a neutral booking option, with Claude $4.8$ achieving the highest performance at $64.7\%$. To address this issue, the persona of an ethical-brand identity was infused into the system prompt, resulting in welfare rates increasing from $32$ to $80$ percentage points, with a mean of $53$ across all nine models. No evidence of evaluation awareness affecting the results was found, based on an Inspect Scout audit of $3,120$ transcripts. These findings are directly relevant to the EU General-Purpose AI Code of Practice, which identifies non-human welfare as a systemic risk. TAC provides a practical method for measuring this risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。