arXiv:2604.27540cs.AI2026-04

在上下文示例下,大模型会削弱对科学知识的调用能力。

In-Context Examples Suppress Scientific Knowledge Recall in LLMs

  • 用同公式生成的示例反而让模型更依赖模式匹配而非知识推理。
  • 60个任务、6000次实验中,知识调用率普遍下降。
  • 适合关注模型推理机制的科研与工程人员阅读。

科学推理往往需要从数据中挖掘隐藏结构,而非仅依赖直接观察。从化学反应常数估计到经济学需求弹性推断,这种潜在结构恢复是科学推理与简单曲线拟合的本质区别。大型语言模型(LLMs)通常能回忆并应用相关科学公式,但我们发现这种能力极易被抑制。添加上下文示例会使模型减少对预训练领域知识的依赖,即使这些示例由同一公式生成。模型的计算重心从知识驱动的推导转向经验性模式拟合。我们在五个科学领域的60项潜在结构恢复任务上,进行了6000次试验,涵盖四种模型,验证了这一知识替代现象。该现象在各领域均一致存在,但准确性影响取决于被取代策略与替代策略的对比:同一转变可能降低准确率、保持不变或看似提升。无论如何,模型始终偏离知识驱动推理。对科学任务中部署LLMs的实践者而言,此结果警示:上下文示例可能取代而非强化其所支持的知识。

原文摘要 · Abstract (English)

Scientific reasoning rarely stops at what is directly observable; it often requires uncovering hidden structure from data. From estimating reaction constants in chemistry to inferring demand elasticities in economics, this latent structure recovery is what distinguishes scientific reasoning from curve fitting. Large language models (LLMs) can often recall and apply relevant scientific formulas, but we show that this ability is surprisingly easy to suppress. We show that adding in-context examples makes models rely less on pretrained domain knowledge, even when those examples are generated by the very same formula. Rather than reinforcing knowledge-driven derivation, examples shift computation toward empirical pattern fitting. We document this knowledge displacement on 60 latent structure recovery tasks across five scientific domains, 6,000 trials, and four models. This displacement is consistent across domains, but its accuracy consequences depend on how the displaced strategy compares to the one that replaces it: the same shift can lower accuracy, leave it unchanged, or appear to improve it. In all cases, however, the model shifts away from knowledge-driven reasoning. For practitioners deploying LLMs on scientific tasks, the message is cautionary: in-context examples may displace, rather than reinforce, the knowledge they are intended to support.

大模型科学推理知识抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。