arXiv:2511.05850cs.IRcs.AI2025-11

Gemini 2.5 Flash在长上下文检索中无中间遗忘问题,事实问答准确率高。

Retrieval Quality at Context Limit

  • 测试不同位置的事实问答,发现模型对中间内容仍能精准召回。
  • 在接近上下文极限时,对关键信息的准确率仍保持高位。
  • 适合需要长文本理解与精准检索的应用场景。

大型语言模型(LLMs)从长上下文中回忆和检索信息的能力对众多实际应用至关重要。先前研究(Liu et al., 2023)指出,当事实位于长上下文中间时,LLMs的检索准确率会显著下降,这一现象称为“中间丢失”(Lost in the Middle, LITM)。我们发现,模型Gemini 2.5 Flash在面对“针中找刺”类问题时,无论文档位置如何,包括接近输入上下文极限时,均能以高准确率回答。结果表明,Gemini 2.5 Flash在简单事实问答任务中不存在“中间丢失”效应,显示出长上下文检索能力的显著提升。

原文摘要 · Abstract (English)

The ability of large language models (LLMs) to recall and retrieve information from long contexts is critical for many real-world applications. Prior work (Liu et al., 2023) reported that LLMs suffer significant drops in retrieval accuracy for facts placed in the middle of large contexts, an effect known as "Lost in the Middle" (LITM). We find the model Gemini 2.5 Flash can answer needle-in-a-haystack questions with great accuracy regardless of document position including when the document is nearly at the input context limit. Our results suggest that the "Lost in the Middle" effect is not present for simple factoid Q\&A in Gemini 2.5 Flash, indicating substantial improvements in long-context retrieval.

长上下文检索准确率Gemini

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。