arXiv:2605.00607cs.CLeess.AS2026-05

用编码探针反向重建模型表示,更准确分析特征贡献。

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

论文配图:Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe
图 1 · 摘自论文原文
  • 反向构建模型表征,用可解释特征还原内部表示
  • 语法和词汇特征独立贡献,说话人特征受训练目标影响大
  • 适合研究模型内部表征机制的研究者使用

探针广泛用于研究语言模型表征中可解码的特征。然而,传统的解码探针方法存在两个局限:不同特征对模型表征的贡献难以直接比较,且特征相关性会影响探针结果。为此,我们提出一种编码探针(Encoding Probe),其反向思路是利用可解释特征来重构模型内部表示。我们在文本与语音变换器模型上评估该方法,使用涵盖声学、音系、句法、词汇及说话人身份的特征集。结果表明,说话人相关效应在不同训练目标和数据集间差异显著,而句法与词汇特征对重建的贡献相互独立。这些发现说明,编码探针为理解模型表征提供了超越可解码性的互补视角。

原文摘要 · Abstract (English)

Probing is widely used to study which features can be decoded from language model representations. However, the common decoding probe approach has two limitations that we aim to solve with our new encoding probe approach: contributions of different features to model representations cannot be directly compared, and feature correlations can affect probing results. We present an Encoding Probe that reverses this direction and reconstructs internal representations of models using interpretable features. We evaluate this method on text and speech transformer models, using feature sets spanning acoustics, phonetics, syntax, lexicon, and speaker identity. Our results suggest that speaker-related effects vary strongly across different training objectives and datasets, while syntactic and lexical features contribute independently to reconstruction. These results show that the Encoding Probe provides a complementary perspective on interpreting model representations beyond decodability.

模型解释编码探针表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。