ECHO让长时任务智能体拥有可审计的记忆,像人一样记住、修正和追溯经历。
ECHO: A Cognitively Inspired, Auditable Memory Plane for Long-Horizon Agents
- 借鉴人类记忆机制设计可审计记忆架构,支持上下文重建与信息修正
- 在1536个问题上达到73.64%的回合召回率,500个长程测试题达88.84%
- 适合需要透明推理路径的AI系统,如医疗、金融等高可信场景
长时序智能体需要能识别相关经验、解决信息修订并提供可核查来源的记忆系统。我们提出ECHO(具身上下文与历史编排),一种受情景编码、巩固、上下文重现、再巩固及执行控制启发的可审计记忆架构与服务原型。该设计为功能类比而非神经等价;实证分析聚焦于信息检索与上下文构建。在1,536个LoCoMo类别1-4问题上,取得96.29%的Hit@10和73.64%的turn Recall@5;在全部500个LongMemEval-S问题上,达到97.60% Hit@10、88.84% turn Recall@5和88.71% session Recall@5。五历史BEAM门失败,且在另500个匹配问答样本中,Mem0 OSS得分为64.84%,而ECHO为41.76%(精确麦内玛尔检验p=0.00107),历史聚类区间跨零。事后审计发现查询扩展规则中存在源特定短语。尽管未引入真实答案字段,但开启扩展的检索得分仍作为描述性开发指标,非独立验证。
原文摘要 · Abstract (English)
Long-horizon agents need memory that identifies relevant experience, resolves revisions, and exposes checkable provenance. We present ECHO (Embodied Context and History Orchestration), an auditable memory architecture and service prototype inspired by episodic encoding, consolidation, contextual reinstatement, reconsolidation, and executive control. This is functional inspiration, not neural equivalence; the empirical analysis focuses on retrieval and context construction. Development runs reach 96.29% Hit@10 and 73.64% turn Recall@5 on 1,536 LoCoMo category 1-4 questions, and 97.60% Hit@10, 88.84% turn Recall@5, and 88.71% session Recall@5 on all 500 LongMemEval-S questions. A five-history BEAM gate fails, and in a separate matched 91-question QA sample Mem0 OSS scores 64.84% versus ECHO's 41.76% (exact McNemar p = 0.00107), with a history-cluster interval crossing zero. A post-hoc audit found source-specific phrases in the query-expansion rules. Although no gold answer field entered the runtime, expansion-enabled retrieval scores are therefore descriptive development measurements, not independent confirmation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。