首个评估大模型长期记忆管理能力的基准,助力优化记忆质量判断。
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
- 构建10种不同记忆模式的评测框架,覆盖8K~128K上下文长度。
- 13个前沿奖励模型表现趋近,新模型普遍优于旧模型。
- 揭示当前奖励模型在长期记忆评估中的优势与局限。
现有研究越来越多地采用以记忆为中心的机制,以分段方式处理长上下文,有效的记忆管理是大语言模型在整个序列中有效传递信息的关键能力。因此,利用奖励模型(RMs)自动且可靠地评估记忆质量至关重要。本文提出MemoryRewardBench,这是首个系统研究奖励模型评估长期记忆管理能力的基准。该基准涵盖长上下文理解与长文本生成任务,包含10种不同记忆管理模式,上下文长度从8K到128K tokens不等。对13个前沿奖励模型的评估显示,开源与专有模型间的性能差距逐渐缩小,新一代模型无论参数量大小均持续优于前代。此外,我们进一步揭示了当前奖励模型在多样化设置下评估大模型记忆管理能力的能力与根本局限。
原文摘要 · Abstract (English)
Existing works increasingly adopt memory-centric mechanisms to process long contexts in a segment manner, and effective memory management is one of the key capabilities that enables large language models to effectively propagate information across the entire sequence. Therefore, leveraging reward models (RMs) to automatically and reliably evaluate memory quality is critical. In this work, we introduce MemoryRewardBench, the first benchmark to systematically study the ability of RMs to evaluate long-term memory management processes. MemoryRewardBench covers both long-context comprehension and long-form generation tasks, featuring 10 distinct settings with different memory management patterns, with context length ranging from 8K to 128K tokens. Evaluations on 13 cutting-edge RMs indicate a diminishing performance gap between open-source and proprietary models, with newer-generation models consistently outperforming their predecessors regardless of parameter count. We further expose the capabilities and fundamental limitations of current RMs in evaluating LLM memory management across diverse settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。