发现语言模型理解话语关系的关键稀疏电路,可跨框架泛化。
Discursive Circuits: How Do Language Models Understand Discourse Relations?
- 提出新任务CuDR,通过最小对比对定位话语理解电路。
- 仅0.2%参数的稀疏电路即可恢复英文话语理解能力。
- 底层处理词汇语义,高层编码话语抽象,通用性强。
Transformer语言模型中哪些组件负责话语理解?我们假设稀疏计算图(称为话语电路)控制模型处理话语关系的方式。与简单任务不同,话语关系涉及更长跨度和复杂推理。为使电路发现可行,我们引入新任务“话语关系下的补全”(CuDR),即在给定话语关系下完成文本。为此构建了专用于激活修补的最小对比对语料库。实验表明,仅约0.2%参数的稀疏电路即可在基于PDTB的CuDR任务中恢复话语理解能力,并能良好泛化至未见过的框架如RST和SDRT。进一步分析显示,低层捕捉词汇语义和共指等语言特征,高层编码话语层级抽象。特征效用在不同框架间保持一致(例如共指支持扩展类关系)。
原文摘要 · Abstract (English)
Which components in transformer language models are responsible for discourse understanding? We hypothesize that sparse computational graphs, termed as discursive circuits, control how models process discourse relations. Unlike simpler tasks, discourse relations involve longer spans and complex reasoning. To make circuit discovery feasible, we introduce a task called Completion under Discourse Relation (CuDR), where a model completes a discourse given a specified relation. To support this task, we construct a corpus of minimal contrastive pairs tailored for activation patching in circuit discovery. Experiments show that sparse circuits ($\approx 0.2\%$ of a full GPT-2 model) recover discourse understanding in the English PDTB-based CuDR task. These circuits generalize well to unseen discourse frameworks such as RST and SDRT. Further analysis shows lower layers capture linguistic features such as lexical semantics and coreference, while upper layers encode discourse-level abstractions. Feature utility is consistent across frameworks (e.g., coreference supports Expansion-like relations).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。