测试100万字长上下文下中文古籍的检索与多跳推理,发现模型表现差异显著。
Retrieval and Multi-Hop Reasoning in 1M-Token Context Windows: Evaluating LLMs on Classical Chinese Text

- 在百万词上下文中实现精准单次信息检索,强模型达100%准确率
- 多跳推理性能随上下文长度呈三种不同衰减模式,512K到1M是关键分水岭
- 实测显示模型能力不能仅凭上下文长度判断,适合评估长文本推理能力
我们评估了五款宣称具备100万词上下文窗口的前沿大模型在古典中文语料上的长上下文检索与推理能力。两项互补实验:测试1在100万词输入中测量单个信息点的检索能力,设置三个深度位置的传记类信息点,并配对真实(训练数据一致)与篡改(训练数据矛盾)版本,以区分真实上下文检索与记忆依赖;测试2旨在探查当检索需中间推理时,长上下文能力是否下降,测量跨三个层级(256K、512K、1M词)的三跳链式推理。结果显示,最强模型(Gemini 3.1 Pro、Claude Opus 4.7、GPT-5.5)在100万词单次检索中均达到100%准确率;但多跳推理表现出三种明显衰退模式:稳定型(Gemini Pro、Claude)在512K内保持80%以上准确率,1M略有下降;悬崖型(GPT-5.5、Qwen3.6-plus)在512K至1M间急剧下滑;平滑衰减型(DeepSeek V4 Pro)则在整个范围内逐渐降低。结果表明,名义上下文长度无法准确反映实际可用的多跳推理能力,而512K至1M之间的过渡是当前旗舰模型的关键分界线。
原文摘要 · Abstract (English)
We evaluate the long-context retrieval and reasoning capabilities of five frontier large language models with advertised 1M-token context windows on a classical Chinese corpus. Two complementary studies are reported. Test 1 measures single-needle retrieval at 1M tokens of input, with three biographical needles planted at three depths and pairs of real (training-prior-consistent) and altered (training-prior-contradicting) variants to separate genuine in-context retrieval from reliance on memorised training data. Test 2, a follow-up designed to probe whether long-context capability degrades when retrieval requires intermediate reasoning, measures three-hop chain traversal across three context tiers (256K, 512K, and 1M tokens). We find that single-needle retrieval at 1M is essentially solved for the strongest models - Gemini 3.1 Pro, Claude Opus 4.7, and GPT-5.5 each achieve 100% - but that multi-hop performance reveals three distinct decay signatures: a stable regime (Gemini Pro, Claude) maintaining greater than 80% accuracy through 512K with modest degradation at 1M; a late-cliff regime (GPT-5.5, Qwen3.6-plus) collapsing sharply between 512K and 1M; and a smooth-decline regime (DeepSeek V4 Pro) decaying gradually across the entire range. The findings suggest that nominal context-window length is a poor proxy for usable long-context multi-hop capability, and that the sharpest discriminator between current 1M-context flagships is the 512K-to-1M transition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。