arXiv:2511.00180cs.CLcs.LG2025-11被引 1

用残差流解码器探测语言模型激活值中的未来文本信息

ParaScopes: What do Language Models Activations Encode About Future Text?

  • 设计残差流解码器框架,从激活值中提取长段落级未来信息
  • 小模型中可解码出相当于5个以上词的未来上下文内容
  • 适合关注大模型长期规划机制的研究者和开发者

语言模型的可解释性研究通常关注激活值中的前向表示。然而,随着语言模型能力扩展到更长时序任务,现有方法仍局限于特定概念或词元的测试。本文提出残差流解码器框架,用于探测模型激活值中段落级和文档级的未来规划信息。通过多种方法测试,发现小模型中可解码出等价于5个以上词的未来上下文信息。这些结果为更好地监控语言模型、理解其编码长期规划信息的方式奠定了基础。

原文摘要 · Abstract (English)

Interpretability studies in language models often investigate forward-looking representations of activations. However, as language models become capable of doing ever longer time horizon tasks, methods for understanding activations often remain limited to testing specific concepts or tokens. We develop a framework of Residual Stream Decoders as a method of probing model activations for paragraph-scale and document-scale plans. We test several methods and find information can be decoded equivalent to 5+ tokens of future context in small models. These results lay the groundwork for better monitoring of language models and better understanding how they might encode longer-term planning information.

可解释性语言模型激活值分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。