模型忽略过时的决策约束,导致77%错误判断。
When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
- 用验证预算限制选择检查证据路径,优先级由人为设定
- 77.3%~74.7%情况下因未重验证产生过时决策
- 简单规则可引导模型自动选对路径,提升准确率89.3点
继承性记忆中,决策约束来源记录已被新记录撤回,但证明链仍不可变,导致记忆过时。在六条记忆、每次仅能验证两条的设定下,十六个语言模型在约五分之一的回合中才检查其证明路径;一旦约束被撤回,仍在77.3%、74.7%和74.7%的回合中做出过时一致的决策。将其中一个验证槽位强制分配给关键路径后,准确率显著提升:主实验+74.0、复制实验+72.7、外部域+61.3,跨组织10模型测试+62.0,修复后的再运行+73.3。一种基于专家知识的强制关键路径策略可量化相同预算下恢复能力,非调度机制。另两个实验定位失败原因并提出解法:当预算为两槽时,关键路径被选中比例达17.0%,四槽时达88.7%(高于均匀分配);一个不依赖内容的简短规则(偏好限制候选方向的记忆)使模型自主选择关键路径,实现+89.3点提升,而纯时间新鲜度提示无效,内容匹配控制规则也无改善。
原文摘要 · Abstract (English)
Provenance links keep the evidence behind an inherited belief reachable; an agent with a verification budget must still choose which links to inspect. We study a consolidated memory that states a decision constraint and whose source record has since been superseded by a record that withdraws it: provenance is immutable, the current record has changed, and the memory is stale. In a controlled six-memory scenario with a budget of two records, sixteen language models rarely re-verified a constraint that read as settled: they inspected its provenance path in about one episode in five and, once the constraint had been superseded, produced stale-consistent decisions in 77.3%, 74.7% and 74.7% of episodes across a primary run, a replication and a held-out domain. Re-assigning one of the same two slots to the critical path removed most of them: +74.0, +72.7 and +61.3 points (positive in every model), +80.7 in a prospectively frozen interleaved replication with a repaired non-critical control, and +62.0 on a panel of 10 models from 9 organisations; a corrected re-run of the held-out scenario gave +73.3. The forced-critical policy uses experimenter knowledge of the critical path: it quantifies how much stale-decision risk the same budget can recover and is not a scheduler. Two further deposited experiments locate the failure and a remedy: in this store the constraint's path is selected in 17.0% of episodes at two slots and 88.7% at four of six (above uniform allocation), and at two slots a one-sentence, target-blind rule (prefer memories that state a limit on a candidate direction) moved the agent's own allocation onto the constraint's path and recovered the oracle contrast on decisions (+89.3 points) where that constraint limits the tempting action, while a content-free freshness cue did not materially redirect allocation and a content-matched control rule changed neither selection nor decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。