探究语言模型何时具备功能性的概念表征能力。
When Do Concepts Become Functionally Sufficient During Language-Model Training?

- 通过注入掩码重构激活,测试各层概念的功能充分性。
- 下游任务保留的软质量远低于重建所需,分布变化小。
- 揭示模型在训练中逐步获得真正有用的内部表示。
深入理解模型及其学习机制,需识别其内部结构何时变得有用,而非仅关注最终状态。本文通过概念动态研究:在每一层和检查点上,分解激活,选取稀疏软掩码,并将掩码重构注入模型。概念分析由此实现功能化检验:掩码仅在干预下能保持目标时才有效。比较了激活重建、线性解码、真实下游保持及学得对齐下的检查点迁移中的充分性。该框架将分解假设视为假设而非可解释性保证,监控跨检查点的功能充分性及源到最终的可重构性。在七种模型的共享固定惩罚操作点上,下游掩码保留的软质量显著低于重建掩码;预测分布漂移仍较小。
原文摘要 · Abstract (English)
Understanding a model and its learning mechanisms in depth requires identifying when its internal structures become useful, rather than simply looking at the final state. We study this through concept dynamics: at each layer and checkpoint, we decompose activations, select sparse soft masks, and inject masked reconstructions into the model. Concept analysis is therefore tested functionally: a mask is useful only insofar as it preserves a target under intervention. We compare sufficiency for activation reconstruction, linear decodability, true downstream preservation, and checkpoint transfer under learned alignment. The framework treats decomposition assumptions as hypotheses rather than interpretability guarantees, monitoring functional sufficiency across checkpoints and source-to-final reconstructability under learned alignment. At the shared fixed-penalty operating point across seven models, downstream masks retain substantially less soft mass than reconstruction masks; predictive-distribution shifts remain small.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。