用生成的隐状态均值池化,能更准确反映语言模型内部状态。
The Truth Lies Somewhere in the Middle (of the Generated Tokens)

- 对生成序列的隐藏状态做均值池化,提取语义表征
- 均值池化比单个词的表示在多个领域表现更优
- 适合研究模型内部行为与表征动态的读者
自回归生成的隐藏状态如何融合为反映语言模型内部状态的表征?尽管生成过程受因果掩码约束,我们发现对所有生成隐状态进行均值池化,其语义表征优于任一单个词。通过核对齐量化其在语言、视觉和蛋白质领域的参考空间表现,均值池化的优势一致表明信息分布在多个生成标记中,而非集中于单一位置。此外,由生成标记得到的表征优于提示标记的表征,且生成过程中的对齐揭示了可解释的模型行为动态。
原文摘要 · Abstract (English)
How should hidden states generated autoregressively be collapsed into a representation that reflects a language model's internal state? Despite tokens being generated under causal masking, we find that mean pooling across their hidden states yields more semantic representations than any individual token alone. We quantify this through kernel alignment to reference spaces in language, vision, and protein domains. The improvement through mean pooling is consistent with information being distributed across generated tokens rather than localized to a single position. Furthermore, representations derived from generated tokens outperform those from prompt tokens, and alignment across generation reveals interpretable dynamics in model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。