arXiv:2504.08202cs.CL2025-04AAAI被引 1

发现长文本模型的内在知识影响随上下文变长而增强,且外在检索与内在回忆能力存在冲突。

Harnessing the Unseen: The Hidden Influence of Intrinsic Knowledge in Long-Context Language Models

  • 提出双能力评估框架,同时考察模型的外部检索与内部记忆能力。
  • Qwen-2.5在长上下文生成中表现优于Llama-3.1,展现更强综合能力。
  • 揭示大模型内在知识潜力常被忽略,需双视角评估性能上限。

近期长上下文语言模型(LCLMs)主要聚焦于利用外部上下文信息,常忽视模型参数化知识的影响。本文首次研究参数化知识对内容生成的作用,发现其影响随上下文长度增加而显著增强。进一步表明,模型的参数化回忆能力(parametric recall ability)与外部检索能力(extrinsic retrieval ability)并非同步提升,甚至更好的外部检索能力会抑制内在回忆能力,限制模型潜力。为此,设计了简单的混合型‘针中寻针’测试,综合评估两种能力。实验结果表明,Qwen-2.5显著优于Llama-3.1,展现出更优的多能力融合潜力;即便更强大的Llama-3.1-70B-Instruct也未表现出更好性能,凸显从双重能力角度评估模型的重要性。

原文摘要 · Abstract (English)

Recent advances in long-context language models (LCLMs), designed to handle extremely long contexts, primarily focus on utilizing external contextual information, often leaving the influence of language models' parametric knowledge underexplored. In this work, we firstly investigate how this parametric knowledge affects content generation and demonstrate that its impact becomes increasingly pronounced as context length extends. Furthermore, we show that the model's ability to utilize parametric knowledge, which we call parametric recall ability, does not improve simultaneously with its ability to leverage contextual knowledge through extrinsic retrieval ability. Moreover, better extrinsic retrieval ability can interfere with the model's parametric recall ability, limiting its full potential. To bridge this gap, we design a simple yet effective Hybrid Needle-in-a-Haystack test that evaluates models based on their capabilities across both abilities, rather than solely emphasizing extrinsic retrieval ability. Our experimental results reveal that Qwen-2.5 models significantly outperform Llama-3.1 models, demonstrating superior potential to combine various abilities. Moreover, even the more powerful Llama-3.1-70B-Instruct model fails to exhibit better performance, highlighting the importance of evaluating models from a dual-ability perspective.

长上下文模型评估知识记忆双能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。