arXiv:2603.15723cs.AI2026-03中稿 · Oral Presentation …

研究大模型在长文本中的问答鲁棒性,发现多跳推理更易受干扰。

Context-Length Robustness in Question Answering Models: A Comparative Empirical Study

  • 控制无关内容长度,测试模型在不同上下文下的表现。
  • 长上下文下准确率下降,多跳任务比单句抽取下降更严重。
  • 适合关注长文档问答与检索增强生成的开发者参考。

大型语言模型越来越多地应用于信息嵌入于长且嘈杂上下文的场景。然而,模型对上下文长度增长的鲁棒性仍不明确。本文通过SQuAD和HotpotQA两个基准,系统增加无关上下文长度,保持答案信号不变,以隔离长度影响。结果表明,随着上下文变长,模型性能持续下降,其中多跳推理任务(HotpotQA)的准确率下降接近单句提取任务(SQuAD)的两倍。这揭示了任务间鲁棒性差异,表明多跳推理尤其易受上下文稀释影响。我们主张在评估模型可靠性时,应显式考察上下文长度鲁棒性,尤其适用于长文档或检索增强生成的应用。

原文摘要 · Abstract (English)

Large language models are increasingly deployed in settings where relevant information is embedded within long and noisy contexts. Despite this, robustness to growing context length remains poorly understood across different question answering tasks. In this work, we present a controlled empirical study of context-length robustness in large language models using two widely used benchmarks: SQuAD and HotpotQA. We evaluate model accuracy as a function of total context length by systematically increasing the amount of irrelevant context while preserving the answer-bearing signal. This allows us to isolate the effect of context length from changes in task difficulty. Our results show a consistent degradation in performance as context length increases, with substantially larger drops observed on multi-hop reasoning tasks compared to single-span extraction tasks. In particular, HotpotQA exhibits nearly twice the accuracy degradation of SQuAD under equivalent context expansions. These findings highlight task-dependent differences in robustness and suggest that multi-hop reasoning is especially vulnerable to context dilution. We argue that context-length robustness should be evaluated explicitly when assessing model reliability, especially for applications involving long documents or retrieval-augmented generation.

问答系统长上下文鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。