高词汇密度会严重压缩大模型有效上下文,影响信息检索
Dense Contexts Are Hard Contexts: Lexical Density Limits Effective Context in LLMs

- 用三组相同长度但信息密度不同的测试,研究词汇密度对模型的影响
- 在高密度场景下,模型检索准确率从接近100%骤降至60%以下
- 揭示真实场景中紧凑高密输入会显著降低模型实际可用上下文
输入长度和相关信息位置常被认为是大语言模型长上下文性能下降的主要原因。本文研究了词汇密度——即上下文中引入新信息的速率——作为第三个被忽视的关键因素,系统性地降低了大模型的有效上下文窗口。我们通过三个“找针”式基准测试,使用9B-685B参数量的开源大模型,在输入长度约12,000词元、针的位置一致的前提下,逐步提高信息密度。结果发现,随着密度上升,模型表现急剧下降:原本在稀疏上下文中近乎完美的模型,在高密度情况下检索得分低于60%。为排除任务类型干扰,我们在每个基准内控制并改变密度,保持其他条件不变。降低密度通常能恢复性能,尤其在高密度区域,降级现象明显缓解。结果表明,有效上下文容量依赖于词汇密度,对实际运行在紧凑高信息密度输入的大模型系统具有直接意义。
原文摘要 · Abstract (English)
Input length and the position of relevant information are widely cited as the primary causes of degraded LLM long-context performance. Here, we study lexical density -- the rate at which a context introduces distinct information -- as a third, largely overlooked factor that systematically reduces the effective context window of LLMs. We quantify the impact of lexical density on open-weight LLMs (9B-685B) using three "find-the-needle" style benchmarks with identical length (~12k tokens) and controlled needle position, but increasing density of information. We observe a sharp performance collapse in higher-density benchmarks: models that are near-perfect in sparse contexts drop below 60% retrieval score on denser ones. To rule out task-type confounds, we vary and control the density within each benchmark while keeping all other properties unchanged. Reducing density generally restores performance, especially in the high-density regimes where degradation appears. These results show that effective context capacity is a function of lexical density, with direct implications for real-world LLM systems operating on compact, information-rich inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。