测试大模型在长文本中推理记忆能力的基准,揭示当前系统严重不足。
RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts

- 构建跨领域长文档基准,评估多跳推理与记忆追踪能力。
- 最强模型仅22.4%准确率,检索与推理均存在明显短板。
- 适合研究长上下文推理、代理记忆与可信AI的学者使用。
大型语言模型及基于LLM的智能体广泛应用于个人助理、企业协作者和自主工作流系统。在这些场景中,记忆(即在长期上下文和多次交互中保留、访问并推理信息的能力)对智能体的可靠性至关重要。本文提出RECON(基于混淆叙事的长上下文推理),一个用于评估复杂推理能力的基准。该基准涵盖三个领域(刑事、医疗、金融)共24份案例文件,每份长度5万至10万词符,测试智能体在六项记忆密集型任务上的表现:重构多跳证据链、传播级联失效、解决来源冲突、反事实推理、满足时间约束以及时间事实检索。现有基准仅检验事实检索或变化检测,而RECON关注变化后的连锁影响——智能体能否追踪哪些下游结论受影响、哪些因独立支持仍成立,以及替代时间线如何演变。评估显示当前架构存在显著缺陷:即使最强的非-Oracle系统也仅达22.4%准确率,检索与推理均成瓶颈。
原文摘要 · Abstract (English)
Large language models and LLM-based agents are widely used as personal chat assistants, enterprise copilots, and autonomous workflow agents. In all these applications, memory (the ability to retain, access, and reason over information accumulated over long contexts and multiple interactions) plays a crucial role in determining the reliability of any agent. We introduce RECON (Reasoning over Extended Contexts with Obfuscated Narratives), a benchmark for evaluating compositional reasoning over long contexts. RECON spans 24 case files across three domains (criminal, medical, and financial), each ranging from 50k to 100k tokens, and tests agents on six memory intensive tasks: reconstructing multi-hop evidence chains, propagating cascading invalidations, resolving source conflicts, counterfactual reasoning, satisfying temporal constraints, and temporal fact retrieval. Recent memory benchmarks evaluate whether agents can retrieve scattered facts or detect if a fact has changed whereas RECON evaluates what happens after the change, whether agents can trace which downstream conclusions are affected, which survive through independent support, and how alternative timelines would have unfolded. Our evaluation reveals substantial limitations across current architectures: even the strongest non-Oracle system reaches only 22.4% Accuracy, with retrieval and reasoning each surfacing as challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。