arXiv:2506.08427cs.CL2025-06ACL被引 1

开源工具Know-MRI可系统解析大模型内部知识机制

Know-MRI: A Knowledge Mechanisms Revealer&Interpreter for Large Language Models

  • 构建可扩展核心模块,自动匹配输入与解释方法
  • 支持多类型输入,统一整合不同解释结果输出
  • 适合研究大模型内部机理的开发者与研究人员

随着大语言模型(LLMs)的持续发展,提升其内部知识机制的可解释性变得日益紧迫。现有解释方法在输入数据格式和输出形式上差异较大,集成工具仅支持特定输入任务,严重限制了实际应用。为此,我们提出开源的Knowledge Mechanisms Revealer&Interpreter(Know-MRI),用于系统分析LLMs中的知识机制。具体而言,我们设计了一个可扩展的核心模块,能自动将不同输入数据与相应解释方法匹配,并统一整合解释输出。用户可根据输入自由选择合适解释方法,从多角度全面诊断模型内部知识机制。代码已公开于https://github.com/nlpkeg/Know-MRI,演示视频见https://youtu.be/NVWZABJ43Bs。

原文摘要 · Abstract (English)

As large language models (LLMs) continue to advance, there is a growing urgency to enhance the interpretability of their internal knowledge mechanisms. Consequently, many interpretation methods have emerged, aiming to unravel the knowledge mechanisms of LLMs from various perspectives. However, current interpretation methods differ in input data formats and interpreting outputs. The tools integrating these methods are only capable of supporting tasks with specific inputs, significantly constraining their practical applications. To address these challenges, we present an open-source Knowledge Mechanisms Revealer&Interpreter (Know-MRI) designed to analyze the knowledge mechanisms within LLMs systematically. Specifically, we have developed an extensible core module that can automatically match different input data with interpretation methods and consolidate the interpreting outputs. It enables users to freely choose appropriate interpretation methods based on the inputs, making it easier to comprehensively diagnose the model's internal knowledge mechanisms from multiple perspectives. Our code is available at https://github.com/nlpkeg/Know-MRI. We also provide a demonstration video on https://youtu.be/NVWZABJ43Bs.

大模型解释知识机制开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。