语言模型的知识并非统一存储,而是随任务变化而分布。
LMs as Task-Specific Knowledge Bases: An Interpretability Analysis

- 知识按任务特定方式编码,不同任务间共享性差
- 同一事实在不同任务中常无法同时激活相关参数
- 适合研究模型可解释性与知识可控性的研究人员
语言模型(LMs)蕴含大量可应用于多种任务的事实性知识,促使人们将其参数视为知识库。知识库的重要特征是:对同一事实的不同查询应返回一致结果,依赖单一真相来源。我们通过行为与机制分析检验语言模型是否满足此特性。结果显示,模型以任务特异性方式编码知识:在某一任务中学到的事实,在其他任务训练过程中常无法共同出现。参数定位实验表明,同一事实在不同任务中由不同的参数子集支撑。此外,链式思维推理的有效性部分源于调用超出评估任务本身的任务特定参数。这些发现表明,模型所知内容与其被提问方式在参数空间中紧密交织,削弱了‘知识库’类比的合理性,并对语言模型中事实性知识的可靠性与可控性带来影响。
原文摘要 · Abstract (English)
Language models (LMs) capture large amounts of factual knowledge applicable to a wide range of tasks, motivating the view of their parameters as a knowledge base. An important property of knowledge bases is that different queries for the same fact return consistent results, drawing on a single source of truth. We investigate whether LMs satisfy this property through behavioral and mechanistic analyses. Our results suggest that they encode knowledge in a task-specific manner. Behaviorally, facts acquired on one task frequently fail to co-emerge on others during training. Parameter localization experiments suggest a mechanistic explanation, revealing distinct parameter subsets underlying different tasks for the same fact. Finally, we show that chain-of-thought reasoning draws part of its effectiveness from engaging task-specific parameters beyond those tied to the evaluation task. Our findings suggest that what the model knows and how it is asked are intertwined in parameter space, undermining the "knowledge base" analogy and carrying implications for the reliability and controllability of factual knowledge in LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。