arXiv:2608.11510q-bio.NCcs.AI2026-08被引 1

研究大模型在冲突任务中的反应机制,发现默认倾向与规则之间存在竞争。

Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task

  • 设计纯语言冲突任务,测试模型对默认颜色响应与规则的对抗。
  • 7个模型中6个表现出显著冲突效应,且在不同条件下激活不同注意力路径。
  • 适合关注大模型认知机制、注意力行为或心理类比研究的读者。

冲突效应在斯特鲁普和侧抑制任务中已有近百年研究,但其机制仍不明确。本文设计了一项仅使用语言的大型语言模型(LLM)冲突任务:提示词引发默认同色完成,显式规则或一致(共轭条件)或冲突(不一致条件)。使用Gemma-2-2B及六款参数从410M到12B的Pythia模型进行实验,结果显示所有模型均表现出强烈的默认同色倾向,七模型中有六显示显著冲突效应。通过因果归因分析、注意力分析与注意力消融,识别出两条不同处理路径:一条是短距离关注表面颜色线索,主要在共轭条件下激活;另一条是长距离关注规则前缀,主要在不一致条件下激活。微调强化默认同色倾向后,不一致表现下降,共轭表现上升;增加规则数量则仅损害不一致表现。这些结果支持一个观点:该任务中的冲突效应源于模型内部权重中的默认映射与上下文规则映射之间的竞争。更广泛而言,本研究展示了大模型作为单一学习网络中默认与规则驱动反应竞争的机制分析工具的潜力。

原文摘要 · Abstract (English)

Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology and neuroscience, but their mechanistic basis is not fully understood. We introduce a verbal-only LLM conflict task in which a prompt stem elicits a default same-color completion and an explicit rule either agrees with (congruent condition) or conflicts with (incongruent condition) the completion. Gemma-2-2B and six Pythia models ranging from 410M to 12B parameters showed strong default same-color tendencies, and six of seven models showed strong congruency effects. Using causal attribution analysis, attention analysis, and attention ablations, we identified distinct processing pathways in these LLMs: a pathway involving short-range attention to a superficial color cue that is preferentially activated in the congruent condition, and a pathway involving long-range attention to the rule prefix that is preferentially activated in the incongruent condition. Fine-tuning that strengthened the default same-color tendency had divergent effects on task conditions, reducing incongruent performance while increasing congruent performance. In contrast, increasing rule set size selectively impaired incongruent performance. These converging findings support an account in which congruency effects in this task arise from competition between an in-weight default mapping and an in-context rule-based mapping. More broadly, our findings illustrate how LLMs can serve as model systems for mechanistic analysis of competition between default and rule-governed response tendencies within a single learned network.

大模型认知注意力机制冲突效应神经机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。