arXiv:2608.04830cs.AI2026-08

评测语言智能体在真实办公流中的记忆能力,发现丰富经验比简洁摘要更助流程延续。

ContextWeave: A Real-World Workflow Benchmark

论文配图:ContextWeave: A Real-World Workflow Benchmark
图 1 · 摘自论文原文
  • 构建1005个可执行任务,模拟14人多月真实办公流,评估记忆对任务表现的影响
  • 最优记忆配置使工作区评分提升至78.20(原68.08),偏好评分从41.50升至70.60
  • 强调记忆需兼顾相关性与执行可靠性,适合研究长期智能体或工作流系统者

随着语言智能体从孤立任务转向长周期、有状态的工作流,内存变得至关重要,但现有评估常将其简化为检索或问答。我们提出ContextWeave,一个纵向基准,用于评估回忆经验是否真正提升智能体在真实办公流中的下游表现。该基准重构了14名参与者经隐私保护处理的多月工作流,形成1,005个可执行任务,其中568个为核心评估任务,包含指令、容器化环境、轨迹及任务专用评分标准。它衡量工作区质量与个体偏好契合度,并辅以相关性、连贯性、可解性及误导性召回的鲁棒性诊断。在固定模型下,最强配置使工作区得分从68.08升至78.20,偏好得分从41.50升至70.60。固定记忆组件时,所有五种基线模型均因召回获益,但增益差异显著。分析表明,可操作的、经验丰富的记忆比紧凑摘要更能支持流程延续并减少冗余探索,但也更易受误导性召回影响。这些发现推动记忆系统不仅优化检索相关性,还需保障执行过程中的可靠性。

原文摘要 · Abstract (English)

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.

智能体记忆评测工作流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。