arXiv:2608.20664cs.AIcs.SE2026-08

测试软件代理在多轮任务中的记忆管理能力,发现仅特定配置能有效留存关键信息。

DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

  • 设计多轮任务依赖早期不可推断证据,用可执行隐式校验评分
  • 唯一有效配置仅达成53.89%正确率,其他均低于12%
  • 适合评估智能体长期记忆机制的科研人员和开发者

DreamBench-SWE 是一个针对软件代理记忆卫生的多会话基准测试,后续任务依赖前序会话中无法推断的证据,并通过可执行的隐藏断言进行评分。原始版本 v2 与独立注册的后续版本 v2.1 审计均完成 360/360 工作单元及 720/720 S3 单元。原始分析显示主对比(DF-hybrid--B5)无显著差异(95/180 vs 89/180;聚类 p=.518,Holm p=1),不支持等效性;C9/C10 仍受 B0 头部空间限制。后续版本中,外部记忆未达 21/180 成功(0.1167 率),确定性逐字事件记忆为 82/180(0.4556),类型+原始参考探针为 83/180(0.4611),一个固定托管的 Mem0 字面存储配置达 97/180(0.5389)。注册的六槽 Family A 因可用槽位缺失而失效(p=1);所有三组与无记忆对比在 Holm 校正后均被拒绝。两组预注册机制对比因预评估合规性拒绝而无效。次级字面存储对逐字对比非确认性且敏感,与参考探针比较亦未拒绝。审计表明 DreamBench-SWE 可作为辨别性可执行评估基准,并刻画出一种精确托管内存配置,但未确立外部系统机制、记忆条件间的优劣、等效性或广泛适用性。原始 v2.0.5 结果与数据保持不变。

原文摘要 · Abstract (English)

DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software tasks depend on non-inferable evidence from earlier sessions and are scored by executable hidden oracles. We report the original scaled v2 fold and a separately preregistered v2.1 successor audit designed after that study but frozen before successor outcome inspection. The successor run completed 360/360 work units and 720/720 S3 cells across four conditions. In the original fold, the primary DF-hybrid--B5 contrast was null (95/180 versus 89/180; clustered p=.518, Holm p=1), not evidence of equivalence, and C9/C10 retained B0-headroom limitations. In the successor, no external memory achieved 21/180 passes (rate 0.1167), deterministic verbatim event memory 82/180 (rate 0.4556), the typed-plus-raw reference probe 83/180 (rate 0.4611), and one pinned hosted Mem0 literal-storage configuration 97/180 (rate 0.5389). The registered six-slot Family A retained unavailable slots at p=1; all three available comparisons against no memory rejected after Holm correction. Both preregistered mechanism contrasts were unavailable after pre-evaluation conformance rejection. The secondary literal-storage-versus-verbatim comparison was nonconfirmatory and sensitivity-dependent, while the comparison with the reference probe did not reject. The audit therefore supports DreamBench-SWE as a discriminating executable profile benchmark and characterizes one exact hosted-memory configuration, but it does not establish an external-system mechanism, superiority among memory-bearing conditions, equivalence, or broad product generality. The original v2.0.5 findings and artifacts remain unchanged.

智能体记忆机制基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。