发现深层模型中间表示的潜在能力,提升迁移学习效果。
Uncovering the Latent Potential of Deep Intermediate Representations

- 提出按层选择最优嵌入的几何方法,避免简单拼接
- 在多种模型和任务中提升性能,深度越大增益越明显
- 揭示语义分布规律,支持跨模态可解释分析
基于海量数据预训练的基础模型在不同深度上形成具有独特语义内容与几何结构的层级嵌入。与仅使用最后一层或浅层混合的普遍做法相反,我们发现任务相关信息在各层间非单调分布,无法通过简单聚合恢复。通过跨多模态的几何与实证研究,表明有效迁移依赖于识别编码任务判别性结构的层及其嵌入的几何组织方式。我们提出层间最优嵌入选择(LOES),一种构造性谱方法,通过在正交与各向同性约束下最小化残差误差,识别任务判别性子空间。为使微调契合此选择原则,进一步提出几何正则化损失(GeoReg),强制类别流形呈现单纯形结构,并稳定微调过程中的表示几何。在多种架构、深度、模态与数据条件下,LOES持续优于标准基线,增益随模型深度增加而提升。除精度外,该方法还揭示了语义因素在各层间的分布,实现跨语言与跨模态的可解释性分析。结果共同表明,层间嵌入几何并非偶然,而是深度模型表征与知识迁移的核心机制。
原文摘要 · Abstract (English)
Foundational Models pretrained on huge amount of data learn representations that evolve across depth, forming a hierarchy of embeddings with distinct semantic content and geometric structure. Contrary to the widespread practice of using only the final layer or shallow mixtures, we show that task-relevant information is distributed non-monotonically across layers and cannot be recovered by naïve aggregation. Through a geometric and empirical study across multiple modalities, we show that effective transfer depends on identifying which layers encode task-discriminative structure and how their embeddings are geometrically organized. We introduce Layer-wise Optimal Embedding Selection (LOES), a constructive spectral method that identifies task-discriminative subspaces by minimizing residual error under orthogonality and isotropy constraints. To align fine-tuning with this selection principle, we further propose Geometric Regularization Loss (GeoReg), which enforces a simplicial structure on class manifolds and stabilizes representation geometry during fine-tuning. Across a wide range of architectures, depths, modalities, and data regimes, LOES consistently outperforms standard baselines, with gains that grow as model depth increases. Beyond accuracy, our method reveals how semantic factors are distributed across layers, thereby enabling cross-lingual and cross-modal interpretability analyses. Together, our results provide strong evidence that layerwise embedding geometry is not incidental but central to how deep models represent and transfer knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。