arXiv:2506.18141cs.CLcs.AI2025-06ACL被引 7

通过稀疏特征共激活,发现大模型中可解释的语义模块。

Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models

  • 用少量提示词捕捉稀疏自编码器特征共激活模式。
  • 删减模块改变输出,增强模块产生反事实结果。
  • 适合研究模型内部机制与可控生成的学者。

我们通过仅使用少量提示词收集的稀疏自编码器(SAE)特征共激活,识别出大型语言模型(LLMs)中语义连贯且上下文一致的网络组件。针对概念-关系预测任务,发现删除这些概念(如国家、词语)和关系(如首都、翻译语言)组件会以可预测方式改变模型输出,而增强这些组件则引发反事实响应。值得注意的是,组合关系与概念组件可生成复合反事实输出。进一步分析表明,大多数概念组件出现在第一层,而更抽象的关系组件集中于后期层。最后,提取的组件比单个特征更全面地捕捉概念与关系,同时保持特异性。总体而言,研究结果暗示知识在模型中具有模块化组织,并推动了高效、精准的LLM操控方法。

原文摘要 · Abstract (English)

We identify semantically coherent, context-consistent network components in large language models (LLMs) using coactivation of sparse autoencoder (SAE) features collected from just a handful of prompts. Focusing on concept-relation prediction tasks, we show that ablating these components for concepts (e.g., countries and words) and relations (e.g., capital city and translation language) changes model outputs in predictable ways, while amplifying these components induces counterfactual responses. Notably, composing relation and concept components yields compound counterfactual outputs. Further analysis reveals that while most concept components emerge from the very first layer, more abstract relation components are concentrated in later layers. Lastly, we show that extracted components more comprehensively capture concepts and relations than individual features while maintaining specificity. Overall, our findings suggest a modular organization of knowledge and advance methods for efficient, targeted LLM manipulation.

大模型机制语义模块反事实生成SAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。