arXiv:2603.17823cs.LGcs.CL2026-03AAAI被引 1

发现大模型内部可解释的功能模块,提升模型透明度。

Discovering Decoupled Functional Modules in Large Language Models

  • 无监督跨层模块发现框架,自动拆解神经元功能
  • 模块语义连贯且对应具体任务专长,性能优于基线
  • 适合研究模型可解释性与模块化设计的学者

理解大语言模型(LLMs)内部功能组织对提升其可信度和性能至关重要。然而,LLMs如何将不同功能组织为模块仍鲜有研究。为此,我们提出功能模块发现问题,并构建无监督大模型跨层模块发现(ULCMOD)框架,同时将整个模型中的大量神经元解耦为功能模块,并识别与这些模块相关的输入样本主题。该框架引入新颖的目标函数和高效的迭代解耦(IterD)算法。大量实验表明,所提方法发现的模块质量高、解耦性强,能捕捉更丰富的语义信息,在多种下游任务中表现优异。定性分析进一步揭示,所发现模块具有语义一致性,对应可解释的功能专长,并在模型中呈现清晰的空间与层次结构。本工作为解析大模型功能模块提供了新工具,填补了大模型可解释性研究的关键空白。

原文摘要 · Abstract (English)

Understanding the internal functional organization of Large Language Models (LLMs) is crucial for improving their trustworthiness and performance. However, how LLMs organize different functions into modules remains highly unexplored. To bridge this gap, we formulate a functional module discovery problem and propose an Unsupervised LLM Cross-layer MOdule Discovery (ULCMOD) framework that simultaneously disentangles the large set of neurons in the entire LLM into modules while discovering the topics of input samples related to these modules. Our framework introduces a novel objective function and an efficient Iterative Decoupling (IterD) algorithm. Extensive experiments show that our method discovers high-quality, disentangled modules that capture more meaningful semantic information and achieve superior performance in various downstream tasks. Moreover, our qualitative analysis reveals that the discovered modules show semantic coherence, correspond to interpretable specializations, and a clear spatial and hierarchical organization within the LLM. Our work provides a novel tool for interpreting the functional modules of LLMs, filling a critical blank in LLM's interpretability research.

可解释性模块发现大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。