LLM通过词语差异性判断构建文本因果结构,像实验一样验证影响。
Words as Difference Makers: How Large Language Models Determine Causal Structure in Text
- 基于词语变化推断因果,利用大量跨领域语料训练差异感知能力
- 自注意力与词嵌入等架构支持系统性差异检测机制
- 为理解LLM如何隐式学习因果提供了实验类逻辑新视角
由于大型语言模型(LLMs)在文本预测上表现卓越,人们认为它们必然具备表征因果与定义结构的‘世界模型’。然而,现代因果推断的主流框架——朱迪亚·珀尔的干预主义方法和奈曼-鲁宾潜在结果框架——难以解释LLMs如何习得因果结构。本文通过论证LLMs采用一种基于差异性逻辑的归纳方法(有时称为变分归纳),解决了这一难题。研究发现,在训练过程中,LLMs需要来自广泛语境的海量文本数据,以识别词序列中的差异制造者与无关因素。此外,本文分析了模型架构特征(如词嵌入和自注意力机制)在实现变分归纳中的作用。这种差异性逻辑本质上与实验方法一致:通过系统性地改变单一条件,确定其对现象的影响。
原文摘要 · Abstract (English)
Because large language models (LLMs) are impressively successful in predicting text, it appears that they must have access to a 'world model' representing causal and definitional structure. However, the dominant formalisms of modern causal inference -- Judea Pearl's interventionist approach and the Neyman-Rubin potential outcomes framework -- struggle to illuminate how LLMs learn causal structure. I resolve this puzzle by arguing that LLMs employ a specific inductive approach based on a difference-making logic -- sometimes called variational induction. I demonstrate how central aspects of this logic are realized during training, where LLMs require enormous amounts of text data from a wide range of contexts to identify difference- and indifference-makers within word sequences. Furthermore, I analyze specific architectural features of LLMs -- such as token embeddings and self-attention -- to determine their roles in variational induction. The difference-making logic of LLMs fundamentally parallels the experimental method, where causal relations are derived by systematically varying individual circumstances to determine their influence on a phenomenon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。