arXiv:2512.15134cs.LGcs.AI2025-12ACL被引 5

检验解释方法能否真正分离不同语义概念,发现看似独立的特征其实相互干扰。

From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?

  • 设计多概念评估框架,同时测试情感、语气、领域等多重概念
  • 多数特征仅对单一概念敏感,但一个概念常分布在多个特征中
  • 操控某特征时往往影响多个概念,说明选择性不足

解释性研究的目标是从未知神经网络激活中恢复出解耦的潜在概念表征。现有方法通常在孤立条件下评估特征质量,并隐含假设各概念相互独立,但这一假设在实际中未必成立。为此,本文提出一种包含情感、领域、语音和时态等概念的多概念评估设置,检验稀疏自编码器(SAEs)和探测器等特征提取方法对各概念的解耦能力。结果显示,特征通常只对单一概念敏感,但同一概念常分布于多个特征中。进一步通过特征操纵实验发现,在理想条件下操控某一特征仍会显著影响多个概念,尽管其交互效应微弱。这表明相关性指标不足以判断操控选择性,仅证明特征位于不同空间也无法保证其仅响应单一概念。该研究强调了多概念评估在解释性研究中的关键作用。

原文摘要 · Abstract (English)

A goal of interpretability is to recover disentangled representations of latent concepts (features) from the activations of neural networks. The quality of features is typically evaluated in isolation, and under implicit independence assumptions that may not hold in practice. Thus, it is unclear to what extent common featurization methods such as sparse autoencoders (SAEs) and probes disentangle one concept from another. We propose a multi-concept evaluation setting using concepts including sentiment, domain, voice, and tense. We evaluate how well featurizers produce disentangled representations of each concept, observing that features are typically sensitive to only one concept, but also that concepts are distributed across many features. Then, we steer these features, measuring whether each concept is independently manipulable, and whether features interact. Even in idealized settings, steering a feature often affects many concepts, despite a near absence of interaction effects. These results suggest that correlational metrics are insufficient to establish steering selectivity, and that demonstrating that two features operate in separate spaces is insufficient to claim that they will be selective for one concept. These results underscore the importance of multi-concept evaluations in interpretability research.

可解释性特征解耦概念操控多概念评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。