arXiv:2605.28831cs.CLcs.AI2026-05被引 1

S3Mem用结构化时空单元高效提取长时问答证据,减少90%以上冗余信息。

S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering

论文配图:S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
图 1 · 摘自论文原文
  • 将视觉、事件、状态等异构历史编码为结构化场景-事件单元
  • 仅用1073/189个证据词(约15.8倍更少)达成领先性能
  • 无需微调或测试时训练,适合处理因果、时序等局部证据

长时问答常需从异构历史中提取稀疏证据,包括事件、物体状态、视觉观测、时间关系和因果步骤。现有记忆接口扩展阅读上下文、检索语义相关片段或暴露图邻域,但未专门设计为在查询时选择紧凑证据。本文提出结构化时空场景-事件记忆(S3Mem),在查询时将文本、视觉和代理使用历史写入结构化场景-事件单元,并路由紧凑证据包给阅读器。其路由器对候选单元、查询锚点及锚点-支持链接进行评分,实现单跳与短多跳证据链选择,无需读者微调或测试时训练。在LoCoMo、EMemBench Visual Games和AMA-Bench上,S3Mem在得分-令牌权衡上表现强劲,尤其在局部事件、状态、时间、因果或来源证据上提升最显著。在LoCoMo上达到0.48 F1和0.40 BLEU,每问题仅使用1,073个证据词,约为参考基线的15.8倍减少;在EMemBench Visual Games上取得最佳F1和第二高准确率,仅需189个词;在AMA-Bench上虽非最高分,但仍具竞争力且使用最少可见证据词。

原文摘要 · Abstract (English)

Long-horizon memory question answering often requires sparse evidence from heterogeneous histories, including events, object states, visual observations, temporal relations, and causal steps. Existing memory interfaces expand reader context, retrieve semantically related chunks, or expose graph neighborhoods, but they are not explicitly designed to select compact evidence for a fixed reader. We propose Structured Spatiotemporal Scene--Event Memory (S3Mem), a query-time memory interface that writes textual, visual, and agent-use histories into structured scene--event units and routes compact evidence packs to the reader. Its router scores candidate units, query anchors, and anchor--support links, enabling both single-hop selection and short multi-hop evidence chains without reader fine-tuning or test-time training. Across LoCoMo, EMemBench Visual Games, and AMA-Bench, S3Mem provides a strong score--token trade-off, with the clearest gains on localized event, state, temporal, causal, or provenance evidence. On LoCoMo, S3Mem reaches \(0.48\) F1 and \(0.40\) BLEU with (1{,}073) evidence tokens per question, about \(15.8\times\) fewer than the LoCoMo reference. On EMemBench Visual Games, it obtains the best F1 and second-best accuracy with only \(189\)tokens.On AMA-Bench, it is not the highest-scoring method, but remains competitive while using the fewest reader-visible evidence tokens.

长时记忆问答系统结构化记忆高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。