arXiv:2602.14044cs.CL2026-02

上下文位置影响大模型事实核查效果,开头结尾证据更有效

Context Shapes LLMs Retrieval-Augmented Fact-Checking Effectiveness

  • 测试不同上下文长度和证据位置对大模型核查的影响
  • 证据放中间时准确率下降,开头或结尾时表现更好
  • 适用于构建更高效的检索增强型事实核查系统

大语言模型在多种任务中展现出强大推理能力,但在长上下文场景下表现不一致。本研究基于HOVER、FEVEROUS和ClimateFEVER三个数据集,评估了五个开源模型(参数量7B、32B、70B,涵盖Llama-3.1、Qwen2.5和Qwen3)在不同上下文长度下的事实核查表现。结果表明,大模型具备显著的参数化事实知识,但随着上下文长度增加,验证准确率普遍下降。与以往研究一致,证据在提示中的位置至关重要:当相关证据位于提示开头或结尾时,准确率显著更高;而放置于中间区域时,性能明显降低。这凸显了提示结构在检索增强型事实核查系统中的关键作用。

原文摘要 · Abstract (English)

Large language models (LLMs) show strong reasoning abilities across diverse tasks, yet their performance on extended contexts remains inconsistent. While prior research has emphasized mid-context degradation in question answering, this study examines the impact of context in LLM-based fact verification. Using three datasets (HOVER, FEVEROUS, and ClimateFEVER) and five open-source models accross different parameters sizes (7B, 32B and 70B parameters) and model families (Llama-3.1, Qwen2.5 and Qwen3), we evaluate both parametric factual knowledge and the impact of evidence placement across varying context lengths. We find that LLMs exhibit non-trivial parametric knowledge of factual claims and that their verification accuracy generally declines as context length increases. Similarly to what has been shown in previous works, in-context evidence placement plays a critical role with accuracy being consistently higher when relevant evidence appears near the beginning or end of the prompt and lower when placed mid-context. These results underscore the importance of prompt structure in retrieval-augmented fact-checking systems.

大模型事实核查提示工程检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。