arXiv:2509.21361cs.CLcs.AI2025-09被引 12

实测发现大模型有效上下文远小于宣称值,问题类型影响显著。

Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs

  • 定义有效上下文窗口,通过标准化测试评估不同长度效果
  • 多数模型在1000词内准确率严重下降,99%未达宣称最大值
  • 问题类型决定有效窗口大小,助降低幻觉提升精度

大型语言模型(LLM)提供商常宣传巨大的最大上下文窗口。为检验其实际应用效果,我们定义了最大有效上下文窗口(MECW),提出针对不同问题类型和窗口大小的测试方法,并建立标准化比较框架以定位模型失效点。我们在多个模型上收集了数十万数据点,发现报告的最大上下文窗口(MCW)与实际有效窗口(MECW)存在显著差异。部分顶尖模型在仅100词上下文时即失效;大多数模型在1000词左右准确率大幅下降。所有模型的实际有效窗口均远低于宣称值,差距可达99%。结果表明,最大有效上下文窗口随任务类型变化,为提升模型准确率、减少幻觉提供了明确可操作的改进方向。

原文摘要 · Abstract (English)

Large language model (LLM) providers boast big numbers for maximum context window sizes. To test the real world use of context windows, we 1) define a concept of maximum effective context window, 2) formulate a testing method of a context window's effectiveness over various sizes and problem types, and 3) create a standardized way to compare model efficacy for increasingly larger context window sizes to find the point of failure. We collected hundreds of thousands of data points across several models and found significant differences between reported Maximum Context Window (MCW) size and Maximum Effective Context Window (MECW) size. Our findings show that the MECW is, not only, drastically different from the MCW but also shifts based on the problem type. A few top of the line models in our test group failed with as little as 100 tokens in context; most had severe degradation in accuracy by 1000 tokens in context. All models fell far short of their Maximum Context Window by as much as 99 percent. Our data reveals the Maximum Effective Context Window shifts based on the type of problem provided, offering clear and actionable insights into how to improve model accuracy and decrease model hallucination rates.

大模型上下文窗口有效性评估幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。