首个评估大模型无意识行为适应能力的基准,揭示其自动学习与规避能力远不如人类。
ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models

- 通过认知科学三类非陈述性记忆设计测试:程序性记忆、启动效应、经典条件反射。
- 17个模型平均表现不足66%,顶尖模型仅达65.3%,显著低于人类水平。
- 适合研究模型自动化行为、具身智能与认知架构改进的学者参考。
现有大语言模型代理的记忆评测仅关注事实的显式回忆,忽视了经验内化为无需主动检索的自动化行为这一隐性记忆机制。该空白至关重要:高效助手应能自动应用已学技能或规避失败动作而无需提示。本文提出ImplicitMemBench,首个系统性评估隐性记忆的基准,基于认知科学中非陈述性记忆的三种标准构念:程序性记忆(干扰后的一次性技能习得)、启动效应(通过配对实验/控制实例引发主题驱动偏差)、经典条件反射(条件刺激与非条件刺激关联影响首次决策)。300项任务采用统一的‘学习-启动-干扰-测试’协议,并以首次尝试得分评估。对17个模型的评估显示严重局限:无一模型超过66%总体得分,表现最佳者为DeepSeek-R1(65.3%)、Qwen3-32B(64.1%)和GPT-5(63.0%),均远低于人类基线。分析发现抑制与偏好存在巨大不对称性(抑制17.6% vs. 偏好75.0%),且普遍存在需架构创新突破的瓶颈。ImplicitMemBench将评估范式从‘代理记住什么’转向‘代理自动执行什么’。
原文摘要 · Abstract (English)
Existing memory benchmarks for LLM agents evaluate explicit recall of facts, yet overlook implicit memory where experience becomes automated behavior without conscious retrieval. This gap is critical: effective assistants must automatically apply learned procedures or avoid failed actions without explicit reminders. We introduce ImplicitMemBench, the first systematic benchmark evaluating implicit memory through three cognitively grounded constructs drawn from standard cognitive-science accounts of non-declarative memory: Procedural Memory (one-shot skill acquisition after interference), Priming (theme-driven bias via paired experimental/control instances), and Classical Conditioning (Conditioned Stimulus--Unconditioned Stimulus (CS--US) associations shaping first decisions). Our 300-item suite employs a unified Learning/Priming-Interfere-Test protocol with first-attempt scoring. Evaluation of 17 models reveals severe limitations: no model exceeds 66% overall, with top performers DeepSeek-R1 (65.3%), Qwen3-32B (64.1%), and GPT-5 (63.0%) far below human baselines. Analysis uncovers dramatic asymmetries (inhibition 17.6% vs. preference 75.0%) and universal bottlenecks requiring architectural innovations beyond parameter scaling. ImplicitMemBench reframes evaluation from "what agents recall" to "what they automatically enact".
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。