用老鼠实验评测大模型,发现其表现远低于真实老鼠。
CheeseBench: Evaluating Large Language Models on Rodent Behavioral Neuroscience Paradigms

- 用文字版迷宫让大模型自主探索,模拟老鼠行为测试。
- 最大模型仅达52.6%成功率,远低于老鼠的78.9%基准。
- 模型性能受界面影响极大,非模型本身能力决定。
我们提出CheeseBench,一个评估大语言模型(LLMs)在九种经典行为神经科学范式(如莫里斯水迷宫、巴恩斯迷宫、T型迷宫等)上的基准,涵盖六种认知维度。每项任务基于同行评审的啮齿类动物实验协议,并提供近似动物表现基准。智能体接收统一系统提示,无任务特定指令,仅通过ASCII文本观测和奖励信号自主发现目标,类似将老鼠置于陌生装置中。我们评估了六款开源大模型(参数量3B至72B),在文本化ASCII渲染环境下,对比随机基线与基于图的强化学习代理。最佳模型Qwen2.5-VL-7B在ASCII输入下平均成功率达52.6%,优于随机基线的32.1%,但低于近似动物基准的78.9%。结果表明:(1)模型规模超过7B后收益递减;(2)更长上下文历史反而降低性能;(3)思维链提示适得其反;(4)视觉-语言架构在7B时有优势,但在32B时反而拖累表现。由于同一模型在不同接口设置下表现从20%到57%不等,说明结果反映的是‘智能体+接口’整体系统,而非模型孤立能力。在统一零样本ASCII协议下,当前开源大模型智能体仍显著低于近似鼠类基准,尤其在空间导航与跨试次状态追踪任务上。
原文摘要 · Abstract (English)
We introduce CheeseBench, a benchmark that evaluates large language models (LLMs) on nine classical behavioral neuroscience paradigms (Morris water maze, Barnes maze, T-maze, radial arm maze, star maze, operant chamber, shuttle box, conditioned place preference, and delayed non-match to sample), spanning six cognitive dimensions. Each task is grounded in peer-reviewed rodent protocols with approximate animal baselines. The agent receives a unified system prompt with no task-specific instructions and must discover goals purely from ASCII text observations and reward signals, much like a rodent placed into an unfamiliar apparatus. We evaluate six open-weight LLMs (3B to 72B parameters) on text-based ASCII renderings and compare against both a random baseline and a graph-based reinforcement learning agent. Our best model (Qwen2.5-VL-7B) reaches 52.6% average success on ASCII input, compared to 32.1% for random agents and 78.9% for approximate rodent baselines. We find that (1) scaling beyond 7B yields diminishing returns, (2) longer context history degrades performance, (3) chain-of-thought prompting hurts rather than helps, and (4) a vision-language architecture provides an advantage at 7B but hurts at 32B. Because the same model's performance ranges from 20% to 57% depending on interface parameters alone, these results characterize the agent-plus-interface system, not the model in isolation. Under this unified zero-shot ASCII protocol, current open-weight LLM agents remain well below approximate rodent reference values, particularly on tasks requiring spatial navigation and within-trial state tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。