arXiv:2506.06485cs.CLcs.AI2025-06ACL被引 10

不同任务对知识依赖度不同,影响大模型在上下文冲突时的表现。

Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict

  • 设计无模型依赖的诊断框架,控制知识源对比冲突
  • 任务需外部知识时,上下文依赖会降低性能
  • 评估时用大模型做裁判可能产生偏差,适合谨慎使用

大型语言模型(LLMs)同时依赖上下文信息与参数化记忆,但二者可能发生冲突。以往研究多聚焦于仅依赖上下文的任务,未明确不同任务对知识利用程度不同时模型的行为。本文提出一种模型无关的诊断框架,在保持底层知识不变的前提下,针对不同知识需求的任务引入可控冲突。在代表性开源与专有模型上的实验表明,冲突下的性能下降由任务特定的知识依赖性与冲突合理性共同驱动;使用推理过程或重复上下文等策略虽能增强上下文依赖,有助于仅需上下文的任务,却会损害需要参数化知识的任务;这些效应导致基于模型的评估存在偏差,质疑了大模型作为评判者可靠性。总体而言,上下文-记忆冲突本质上是任务相关的,呼吁在部署与评估中采用任务感知的平衡策略。

原文摘要 · Abstract (English)

Large language models (LLMs) draw on both contextual information and parametric memory, yet these sources can conflict. Prior studies have largely examined this issue in contextual question answering, implicitly assuming that tasks should rely on the provided context, leaving unclear how LLMs behave when tasks require different types and degrees of knowledge utilization. We address this gap with a model-agnostic diagnostic framework that holds underlying knowledge constant while introducing controlled conflicts across tasks with varying knowledge demands. Experiments on representative open-weight and proprietary LLMs show that performance degradation under conflict is driven by both task-specific knowledge reliance and conflict plausibility; that strategies such as rationales or context reiteration increase context reliance, helping context-only tasks but harming those requiring parametric knowledge; and that these effects bias model-based evaluation, calling into question the reliability of LLMs as judges. Overall, our findings reveal that context-memory conflict is inherently task-dependent and motivate task-aware approaches to balancing context and memory in LLM deployment and evaluation.

大模型知识冲突评估偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。