评测大模型在俄语对话中的长期记忆能力,揭示不同记忆机制的优劣。
RUMBA: Russian User Memory Benchmark

- 构建俄语对话记忆基准,细分类别并融合时间与上下文推理。
- 包含带时间戳的多轮对话,需跨会话检索与组合推理。
- 可用于诊断模型缺陷,适合研究长程记忆的学者使用。
大模型处理长期记忆的能力日益重要,但现有基准仍以英语为主,依赖聚合检索指标,无法捕捉长程上下文、时间信息与推理之间的交互。为此,我们提出RUMBA(俄语用户记忆基准)——一个面向长期对话记忆的新基准,提供细粒度的记忆相关问题分类,并采用统一方法论,综合考虑语义类型、会话范围、时间推理及时间表达显式性。RUMBA包含带时间戳的用户-助手对话与问答对,要求模型在跨会话情境下进行检索、整合与推理。尽管专为俄语设计,我们也基于相同方法提供了对齐的英文子集。我们评估了当前主流记忆系统与长上下文模型,证明RUMBA可作为诊断工具,分析模型在不同测试片段上的表现,识别各类记忆机制的优势与失效模式。
原文摘要 · Abstract (English)
The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate retrieval metrics, failing to capture interactions between long-range context, temporal information, and reasoning. To address this, we introduce RUMBA (Russian User Memory BenchmArk) - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric question types and a unified methodology accounting for semantic type, session scope, temporal reasoning, and the explicitness of temporal expressions. RUMBA consists of timestamped user-assistant dialogues with QA pairs requiring retrieval, combination, and reasoning across sessions. While designed for Russian, we also provide an aligned English subset under the same methodology. We evaluate contemporary memory systems and long-context models, and show how RUMBA serves as a diagnostic tool to analyze model behavior across benchmark slices and identify strengths and failure modes of different memory mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。