不同架构模型的表征对齐受结构约束影响,提升跨模型特征兼容性。
Cross-Model Semantics in Representation Learning
- 通过线性调节与修正路径设计,增强模型间表征一致性
- 结构正则化使表征几何在架构变化下更稳定
- 适用于模型蒸馏与模块化学习系统设计
深度网络内部表征通常对架构特定选择敏感,引发关于学习结构在模型间稳定性、对齐性和可迁移性的疑问。本文研究线性塑造算子和修正路径等结构约束如何影响不同架构间内部表征的兼容性。基于先前关于结构变换与收敛性的发现,我们构建了一个测量和分析具有不同但相关架构先验的网络之间表征对齐的框架。结合理论分析、实证探测与受控迁移实验,我们证明结构规律会诱导出在架构变化下更稳定的表征几何。这表明某些归纳偏置不仅支持单个模型内的泛化,还能提升跨模型学习特征的互操作性。最后讨论了表征可迁移性在模型蒸馏、模块化学习及鲁棒学习系统原理性设计中的意义。
原文摘要 · Abstract (English)
The internal representations learned by deep networks are often sensitive to architecture-specific choices, raising questions about the stability, alignment, and transferability of learned structure across models. In this paper, we investigate how structural constraints--such as linear shaping operators and corrective paths--affect the compatibility of internal representations across different architectures. Building on the insights from prior studies on structured transformations and convergence, we develop a framework for measuring and analyzing representational alignment across networks with distinct but related architectural priors. Through a combination of theoretical insights, empirical probes, and controlled transfer experiments, we demonstrate that structural regularities induce representational geometry that is more stable under architectural variation. This suggests that certain forms of inductive bias not only support generalization within a model, but also improve the interoperability of learned features across models. We conclude with a discussion on the implications of representational transferability for model distillation, modular learning, and the principled design of robust learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。