长上下文模型在简单检索任务中表现差,需足够推理步骤才能改进。
Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
- 通过思维链提示引导多步推理提升检索能力
- 少量推理步骤下模型仍会失败,需足够步骤才有效
- 适合关注长上下文推理机制的研究者
长上下文语言模型(LCLMs)因具备大容量上下文窗口而日益流行。然而,尽管在标准长上下文检索任务中表现近乎完美,我们的评估显示其在某些基础任务上仍会失败。后续发现,只要提供足够数量的推理步骤,并辅以特定思维链(CoT)提示,即可显著改善表现。这一结果表明,解决特定长上下文任务可能需要采用长思维链方法;而以往的长上下文基准测试通常忽视长推理的必要性,将任务简单视为直接问答。本研究强调了在长上下文场景中引入深度推理的重要性。
原文摘要 · Abstract (English)
Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long-context retrieval tasks, our evaluations demonstrate they fail in some basic cases. Later, we find they can be well addressed with a sufficient number of reasoning steps, guided by specific CoT prompts. This result emphasizes the potential necessity of solving specific long-context tasks using long-CoT methods, while previous long-context benchmarks always ignore the necessity of long reasoning for long-context tasks and treat them as direct QA tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。