LLM错误常受无关上下文误导,呈现类别的错泛化行为。
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
- 模型基于抽象类别与上下文特征结合推断答案
- 39种事实回忆任务中,下层构建类别表征,上层生成答案
- 两类推理路径竞争决定输出,适合研究模型偏差者阅读
大语言模型在自然语言任务中的成功伴随担忧:它们是否仅是机械复现预训练数据的‘随机鹦鹉’?本文研究无关上下文幻觉现象,发现模型错误源于一种结构化但缺陷明显的机制——类别式(误)泛化。通过行为分析和对Llama-3、Mistral、Pythia在39种事实召回关系类型上的可解释性实验,揭示模型内部计算中:(i) 抽象类别表征在低层形成,高层逐步细化为具体答案;(ii) 特征选择由两个竞争回路控制——一个依赖查询本身,另一个吸收上下文线索,其相对影响力决定最终输出。该发现表明,大模型虽能通过形式训练提取抽象规律,但泛化依赖不可靠的上下文提示,我们称之为‘随机变色龙’。
原文摘要 · Abstract (English)
The widespread success of large language models (LLMs) on NLP benchmarks has been accompanied by concerns that LLMs function primarily as stochastic parrots that reproduce texts similar to what they saw during pre-training, often erroneously. But what is the nature of their errors, and do these errors exhibit any regularities? In this work, we examine irrelevant context hallucinations, in which models integrate misleading contextual cues into their predictions. Through behavioral analysis, we show that these errors result from a structured yet flawed mechanism that we term class-based (mis)generalization, in which models combine abstract class cues with features extracted from the query or context to derive answers. Furthermore, mechanistic interpretability experiments on Llama-3, Mistral, and Pythia across 39 factual recall relation types reveal that this behavior is reflected in the model's internal computations: (i) abstract class representations are constructed in lower layers before being refined into specific answers in higher layers, (ii) feature selection is governed by two competing circuits -- one prioritizing direct query-based reasoning, the other incorporating contextual cues -- whose relative influences determine the final output. Our findings provide a more nuanced perspective on the stochastic parrot argument: through form-based training, LLMs can exhibit generalization leveraging abstractions, albeit in unreliable ways based on contextual cues -- what we term stochastic chameleons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。