arXiv:2506.00069cs.CLcs.AI2025-06被引 9

测试大模型在长对话中对上下文的敏感性,发现性能可下降73%。

Evaluating the Sensitivity of LLMs to Prior Context

  • 设计新基准,系统测试多轮对话中上下文长度与类型的影响
  • 部分模型多轮问答准确率下降最高达73%,大模型也降32%
  • 任务描述位置优化可提升准确率3.5倍,适合部署调优参考

随着大语言模型(LLMs)在多轮对话等持续交互场景中的广泛应用,理解长上下文对其性能的影响至关重要。现有主流基准主要聚焦单轮问答(QA),难以捕捉多轮交互效应。为此,我们引入一组新型基准,系统性地改变先验上下文的量与性质。我们在多个典型LLM(包括GPT、Claude、Gemini)上评估其在这些基准上的表现,以衡量对上下文变化的敏感度。结果表明,多轮交互中,某些模型在多项选择题上的性能可大幅下降,最高达73%;即使是高能力模型GPT-4o,准确率也最多下降32%。值得注意的是,大模型与小模型的相对表现并非始终可预测。此外,合理安排任务描述在上下文中的位置,能显著缓解性能下降,使准确率提升高达3.5倍。这些发现凸显了在模型设计、评估与优化中应对上下文敏感性的必要性。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed in multi-turn dialogue and other sustained interactive scenarios, it is essential to understand how extended context affects their performance. Popular benchmarks, focusing primarily on single-turn question answering (QA) tasks, fail to capture the effects of multi-turn exchanges. To address this gap, we introduce a novel set of benchmarks that systematically vary the volume and nature of prior context. We evaluate multiple conventional LLMs, including GPT, Claude, and Gemini, across these benchmarks to measure their sensitivity to contextual variations. Our findings reveal that LLM performance on multiple-choice questions can degrade dramatically in multi-turn interactions, with performance drops as large as 73% for certain models. Even highly capable models such as GPT-4o exhibit up to a 32% decrease in accuracy. Notably, the relative performance of larger versus smaller models is not always predictable. Moreover, the strategic placement of the task description within the context can substantially mitigate performance drops, improving the accuracy by as much as a factor of 3.5. These findings underscore the need for robust strategies to design, evaluate, and mitigate context-related sensitivity in LLMs.

大模型上下文敏感多轮对话性能评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。