提出新评估框架,区分大模型代理的判断失误与执行失败。
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
- 设计可计算公平参考策略的供应链补货基准,分离状态感知与行动控制。
- 四款主流模型诊断率84%-88%,但技能得分从-0.23到0.62差异显著。
- 发现执行失败有双重表现:响应不足或响应过度,均影响最终结果。
大模型代理在多周决策任务中面临状态不可见的问题,导致最终成本无法揭示失败原因——可能是误判世界,也可能是正确判断后仍未能行动(知行差距)。现有评估无法区分这两类失败,因参考策略要么依赖代理无法获取的特权信息,要么完全缺失。本文提出STOCKTAKE,一个26周的供应链补货基准,基于六种隐藏因子的分解部分可观马尔可夫决策过程,使公平参考策略可计算:每个因子对应一个精确贝叶斯滤波器,使用与代理相同的观测流生成推演策略。将每轮评分介于症状盲基线(0)与该理想参考(1)之间,获得技能得分;同时对每周推理文本进行分析,得到陈述信念检测延迟和知行比率,实现状态估计与控制的独立测量。在50个种子下具有定制压力特征的实验中,Claude Sonnet 5、GPT-5.4、DeepSeek-V4-Pro和Grok 4.5均能检测到84%-88%的隐藏故障,通常在故障发生后一周内识别,但技能得分范围为0.62至-0.23:其中两款低于症状盲基线,却比另外两款更快命名因素。当压力持续时,所有模型在正确诊断的压力周中仍有34%-43%发生缺货,此率部分反映其察觉到的周度严重性。该率还与技能呈反向关系:低于基线的两款模型在诊断周中缺货最少,说明响应不足只是问题的一面,另一面是响应成本超过保护价值。STOCKTAKE量化了这一双重失败。
原文摘要 · Abstract (English)
LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read it correctly and still failed to act (the knowing-doing gap). Existing evaluations cannot separate these two failures; their reference policies either read privileged information the agent never sees, or are missing altogether. We introduce STOCKTAKE, a 26-week supply-chain replenishment benchmark built as a factored partially observable Markov decision process with six hidden factor processes, designed so that a fair reference policy is computable: an exact Bayes filter per factor drives a rollout policy on the identical observation stream the agent receives. Scoring each run between a symptom-blind base-stock floor (0) and this oracle (1) yields a skill score, and grading each week's written rationale yields a stated-belief detection lag and a knowing-doing rate, so state estimation and control are measured separately. On fifty seeds with curated stress profiles, Claude Sonnet 5, GPT-5.4, DeepSeek-V4-Pro, and Grok 4.5 detect 84-88% of hidden failures, typically within a week of onset, yet span skill scores from 0.62 to -0.23: two of the four end below the symptom-blind floor while naming factors slightly faster than the two that beat it. The failure has two faces. Where stress persists, 34-43% of correctly diagnosed stress weeks still end in stockout for every model, a rate that partly reflects the severity of the weeks models notice. That rate also runs opposite to skill: the two models under the floor stock out least on diagnosed weeks, so under-response is only one face of the gap, and their traces point to the other, responses whose cost exceeds what they protect. STOCKTAKE measures both directions of that failure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。