测试软件代理在多轮任务中的记忆管理能力,发现仅特定配置能有效留存关键信息。
DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents
- 设计多轮任务依赖早期不可推断证据,用可执行隐式校验评分
- 唯一有效配置仅达成53.89%正确率,其他均低于12%
- 适合评估智能体长期记忆机制的科研人员和开发者
DreamBench-SWE 是一个针对软件代理记忆卫生的多会话基准测试,后续任务依赖前序会话中无法推断的证据,并通过可执行的隐藏断言进行评分。原始版本 v2 与独立注册的后续版本 v2.1 审计均完成 360/360 工作单元及 720/720 S3 单元。原始分析显示主对比(DF-hybrid--B5)无显著差异(95/180 vs 89/180;聚类 p=.518,Holm p=1),不支持等效性;C9/C10 仍受 B0 头部空间限制。后续版本中,外部记忆未达 21/180 成功(0.1167 率),确定性逐字事件记忆为 82/180(0.4556),类型+原始参考探针为 83/180(0.4611),一个固定托管的 Mem0 字面存储配置达 97/180(0.5389)。注册的六槽 Family A 因可用槽位缺失而失效(p=1);所有三组与无记忆对比在 Holm 校正后均被拒绝。两组预注册机制对比因预评估合规性拒绝而无效。次级字面存储对逐字对比非确认性且敏感,与参考探针比较亦未拒绝。审计表明 DreamBench-SWE 可作为辨别性可执行评估基准,并刻画出一种精确托管内存配置,但未确立外部系统机制、记忆条件间的优劣、等效性或广泛适用性。原始 v2.0.5 结果与数据保持不变。
原文摘要 · Abstract (English)
DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software tasks depend on non-inferable evidence from earlier sessions and are scored by executable hidden oracles. We report the original scaled v2 fold and a separately preregistered v2.1 successor audit designed after that study but frozen before successor outcome inspection. The successor run completed 360/360 work units and 720/720 S3 cells across four conditions. In the original fold, the primary DF-hybrid--B5 contrast was null (95/180 versus 89/180; clustered p=.518, Holm p=1), not evidence of equivalence, and C9/C10 retained B0-headroom limitations. In the successor, no external memory achieved 21/180 passes (rate 0.1167), deterministic verbatim event memory 82/180 (rate 0.4556), the typed-plus-raw reference probe 83/180 (rate 0.4611), and one pinned hosted Mem0 literal-storage configuration 97/180 (rate 0.5389). The registered six-slot Family A retained unavailable slots at p=1; all three available comparisons against no memory rejected after Holm correction. Both preregistered mechanism contrasts were unavailable after pre-evaluation conformance rejection. The secondary literal-storage-versus-verbatim comparison was nonconfirmatory and sensitivity-dependent, while the comparison with the reference probe did not reject. The audit therefore supports DreamBench-SWE as a discriminating executable profile benchmark and characterizes one exact hosted-memory configuration, but it does not establish an external-system mechanism, superiority among memory-bearing conditions, equivalence, or broad product generality. The original v2.0.5 findings and artifacts remain unchanged.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。