arXiv:2606.01223cs.CLcs.AI2026-06

构建对话长时记忆的反思能力评测基准,推动模型从记忆检索迈向深层理解。

Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue

论文配图:Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue
图 1 · 摘自论文原文
  • 提出分层框架REMIND,通过逐步推理实现从证据感知到抽象理解的建构。
  • 在2.6万条标注问答上验证,现有模型准确率不足40%,新方法显著提升性能。
  • 适合研究长对话理解、具身认知与多模态推理的学者参考。

尽管长上下文建模取得进展,现有评测仍局限于显式事实记忆的召回,无法衡量将分散、多模态线索整合为高层解释所需的反思记忆能力。为此,我们提出RefMem-Bench,一个面向长时对话中反思记忆的评测基准,包含26,000条标注问答实例,涵盖八类反思记忆维度和三种任务形式,要求模型超越表层检索,从交互历史中分布式的证据推断潜在含义。为增强反思记忆能力,我们提出反射记忆诱导框架REMIND,将反思记忆视为渐进意义建构过程,结合问题条件化证据检索、显著性感知定位与抽象层级监督,并采用渐进式反射对齐,将高层推理能力蒸馏至事实推理路径。实验表明,RefMem-Bench对当前模型构成重大挑战,而REMIND通过逐步证据感知、定位与抽象,在答案准确率与记忆召回率上均持续提升。

原文摘要 · Abstract (English)

Despite substantial progress in long-context modeling, existing benchmarks remain confined to factual memory for explicit recall, failing to measure the reflective memory required to synthesize fragmented, multimodal cues into high-level interpretations. To address this gap, we introduce RefMem-Bench, a benchmark for reflective memory in long-horizon dialogue. RefMem-Bench contains 26K annotated QA instances with eight reflective-memory dimensions and three task formats, requiring models to move beyond surface-level retrieval and infer latent meanings from evidence distributed across interaction histories. To enhance reflective memory capability, we propose REflective Memory INDuction (REMIND), a hierarchical framework that treats reflective memory as progressive meaning construction. REMIND couples question-conditioned evidence retrieval, salience-aware grounding, and abstraction-level supervision, and uses Progressive Reflective Alignment to distill high-level reflective reasoning into the factual inference pathway. Experiments show RefMem-Bench poses a substantial challenge to current models, while REMIND consistently improves both answer accuracy and memory recall through progressive evidence perception, grounding, and abstraction.

长对话反思记忆评测基准多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。