arXiv:2509.13316cs.CLcs.LG2025-09被引 5

揭示激活值语义化方法可能只是在复述输入,而非暴露模型内部机理。

Do Activation Verbalization Methods Convey Privileged Information?

  • 用第二语言模型将激活值转为自然语言描述
  • 实验发现语义化内容多来自生成器模型的知识,而非目标模型
  • 现有评测数据集无法真实反映方法对模型内部的理解能力

近期可解释性方法提出使用第二个语言模型(verbalizer LLM)将大语言模型(LLM)的内部表示转化为自然语言描述,旨在揭示目标模型如何表征和处理输入。但这些激活值语义化方法是否真能提供关于目标模型内部运作的特权信息,还是仅传递了输入内容的信息?我们系统评估了先前工作中常用的方法与数据集,发现无需访问目标模型内部即可在这些基准上取得良好表现,表明现有数据集不适合作为评测标准。进一步控制实验显示,生成的语义化内容往往反映的是生成器模型自身的参数化知识,而非目标模型的内在知识。结果表明,亟需设计更严谨的基准测试与实验控制,以真正评估语义化方法是否能揭示大语言模型的真实运行机制。

原文摘要 · Abstract (English)

Recent interpretability methods have proposed to translate LLM internal representations into natural language descriptions using a second verbalizer LLM. This is intended to illuminate how the target model represents and operates on inputs. But do such activation verbalization approaches actually provide privileged knowledge about the internal workings of the target model, or do they merely convey information about the inputs provided to it? We critically evaluate popular verbalization methods and datasets used in prior work and find that one can perform well on such benchmarks without access to target model internals, suggesting that these datasets are not ideal for evaluating verbalization methods. We then run controlled experiments which reveal that verbalizations often reflect the parametric knowledge of the verbalizer LLM that generated them, rather than the knowledge of the target LLM whose activations are decoded. Taken together, our results indicate a need for targeted benchmarks and experimental controls to rigorously assess whether verbalization methods provide meaningful insights into the operations of LLMs.

可解释性语言模型激活分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。