arXiv:2510.14936cs.LGcs.AI2025-10被引 4

提出新方法,直接从权重解析神经网络内部机制。

Circuit Insights: Towards Interpretability Beyond Activations

  • 从权重直接分析特征,无需依赖解释模型或数据集。
  • 发现激活背后的组件交互,揭示传统方法无法捕捉的电路动态。
  • 适合研究模型可解释性与机制分析的科研人员。

可解释AI与机械可解释性旨在揭示神经网络的内部结构,其中电路发现是理解模型计算的核心工具。现有方法依赖人工检查,仅适用于简单任务。自动化可解释性虽提升可扩展性,但常忽略特征间交互,且严重依赖外部大模型和数据质量。近期提出的转换器(Transcoders)可将特征归因分解为输入相关与输入无关成分,为系统化电路分析奠定基础。基于此,我们提出WeightLens与CircuitLens两种互补方法,突破激活依赖分析。WeightLens直接从学习到的权重解析特征,无需解释模型或数据集,在不依赖上下文的特征上性能达到或超过现有方法。CircuitLens揭示特征激活如何由组件间交互产生,识别出激活分析无法察觉的电路级动态。两者结合显著提升可解释性鲁棒性,增强可扩展的机械分析能力,同时保持高效与高质量。

原文摘要 · Abstract (English)

The fields of explainable AI and mechanistic interpretability aim to uncover the internal structure of neural networks, with circuit discovery as a central tool for understanding model computations. Existing approaches, however, rely on manual inspection and remain limited to toy tasks. Automated interpretability offers scalability by analyzing isolated features and their activations, but it often misses interactions between features and depends strongly on external LLMs and dataset quality. Transcoders have recently made it possible to separate feature attributions into input-dependent and input-invariant components, providing a foundation for more systematic circuit analysis. Building on this, we propose WeightLens and CircuitLens, two complementary methods that go beyond activation-based analysis. WeightLens interprets features directly from their learned weights, removing the need for explainer models or datasets while matching or exceeding the performance of existing methods on context-independent features. CircuitLens captures how feature activations arise from interactions between components, revealing circuit-level dynamics that activation-only approaches cannot identify. Together, these methods increase interpretability robustness and enhance scalable mechanistic analysis of circuits while maintaining efficiency and quality.

可解释性权重分析电路发现机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。