arXiv:2502.17169cs.CL2025-02Conference of the …被引 2

用复杂逻辑文本测试大模型真实长上下文推理能力

Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)

  • 构建含2048个逻辑命题的长篇文本,模拟真实推理场景
  • 实测有效上下文窗口仅约128个命题,远低于宣称能力
  • 适合评估大模型在复杂任务中真正的长程推理水平

大型语言模型展现出令人期待的长上下文处理能力,近期模型声称可支持接近百万词元的上下文窗口。然而,支撑这些说法的评估常采用简单检索任务或用无关文本填充的合成任务,模型可能轻易识别并丢弃。本文生成长达25,000 GPT-4词元的简化英语文本,包含最多2048个一阶逻辑命题,并设计矛盾检测中的证据检索任务。文本中充斥难以区分的干扰项,且被证明不会影响真实证据。实验表明,在真实干扰下,有效上下文窗口显著缩小,早在128个命题处即开始崩溃。

原文摘要 · Abstract (English)

Large language models demonstrate promising long context processing capabilities, with recent models touting context windows close to one million tokens. However, the evaluations supporting these claims often involve simple retrieval tasks or synthetic tasks padded with irrelevant text, which the models may easily detect and discard. In this work, we generate lengthy simplified English text with first-order logic representations spanning up to 2048 clauses (around 25k GPT-4 tokens). We formulate an evaluation task with evidence retrieval for contradiction detection. The long, homogeneous text is filled with distractors that are both hard to distinguish from relevant evidences and provably not interfering with them. Our evaluation of evidence retrieval shows that the effective context window is much smaller with realistic distractors, already crumbling at 128 clauses.

大模型评测长上下文逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。