LLM在复杂推理中易出现记忆漂移,远早于现有测试揭示的长度。
Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- 用图结构推理任务替代简单续写,更真实评估模型能力。
- 实际有效上下文长度比基准测试显示短得多,记忆漂移提前发生。
- 连专用推理模型o1也难逃早期遗忘,适合长程推理研究者关注。
近期评估基准试图刻画大语言模型(LLMs)的有效上下文长度与遗忘倾向,但多依赖简单的'大海捞针'式检索或续写任务,难以反映模型在信息密集场景下的真实表现。因此,我们主张在更复杂的推理任务中评估模型,要求其从文本中推断出结构化关系知识——如从潜在噪声的自然语言内容中构建图谱。尽管输入文本可视为图生成的结果,其结构未显式表达,关联需从分散的文本线索中推断,这些线索被长距离上下文分隔并夹杂无关信息。我们的发现表明,当执行此类关系推理时,LLMs表现出的记忆漂移和上下文遗忘现象发生在远短于现有基准所提示的有效长度。基于此,我们提出了主流LLMs在复杂推理任务中的最优使用建议。此外,我们还发现即使专为推理设计的模型如OpenAI o1,在此类设置下仍易出现早期记忆漂移。这些结果揭示了模型从非结构化输入中抽象结构化知识的能力存在显著局限,并凸显了需通过架构改进以增强长程推理能力。
原文摘要 · Abstract (English)
Recently proposed evaluation benchmarks aim to characterize the effective context length and the forgetting tendencies of large language models (LLMs). However, these benchmarks often rely on simplistic 'needle in a haystack' retrieval or continuation tasks that may not accurately reflect the performance of these models in information-dense scenarios. Thus, rather than simple next token prediction, we argue for evaluating these models on more complex reasoning tasks that requires them to induce structured relational knowledge from the text - such as graphs from potentially noisy natural language content. While the input text can be viewed as generated in terms of a graph, its structure is not made explicit and connections must be induced from distributed textual cues, separated by long contexts and interspersed with irrelevant information. Our findings reveal that LLMs begin to exhibit memory drift and contextual forgetting at much shorter effective lengths when tasked with this form of relational reasoning, compared to what existing benchmarks suggest. With these findings, we offer recommendations for the optimal use of popular LLMs for complex reasoning tasks. We further show that even models specialized for reasoning, such as OpenAI o1, remain vulnerable to early memory drift in these settings. These results point to significant limitations in the models' ability to abstract structured knowledge from unstructured input and highlight the need for architectural adaptations to improve long-range reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。