不同视觉模型训练后竟共享同一几何结构,可跨领域迁移且抗校准干扰。
The Cross-Architecture Substrate: A Domain-Transcendent, Calibration-Surviving Geometric Invariant of Modern Vision Encoders

- 发现13种视觉模型的前16主方向收敛为同一几何体,称跨架构底座。
- 在4个视觉域上中位匹配度达0.679,跨8域仍超0.604,且不受校准影响。
- 可用于无标签迁移评估、低资源探测、知识蒸馏等场景,性能超越现有方法。
不同视觉神经网络(用于分类、对比、重建或图文匹配)应有不同内部表征,但实验显示它们并无差异。训练后,13种现代视觉编码器的前16个主变分方向收敛至同一16维几何对象,称为跨架构底座。通过主成分分析(PCA)、中心核对齐(CKA)和Pang 2026校准法研究该结构。其在4个视觉域(自然照片、医学CT、卫星、显微镜)间中位Procrustes-CKA为0.679,在8个域(新增草图、深度、热红外、天文)中为0.604,所有配对均>0.40。该结构全局(7.4倍判别力vs-MAE分离,n=13,394)与局部(4.82–5.30,p<10⁻⁴⁴)均抗校准。它非像素统计(0.263)、非Gabor特征(0.31)、非随机投影(0.041),且在训练前10%即出现,而准确率持续上升。提出四项应用:无标签迁移性筛选(比LogME快3倍,+0.15 Kendall-tau);四类域检测(准确率99.6%);冻结低样本探针(16维胜过768维DINOv2,N=50时提升3.78个百分点);教师无关蒸馏辅助匹配(33对中峰值增益7.56点,标签比例10%)。底座不跨模态,不提升跨范式蒸馏,也无法预测迁移质量(ρ=0.08)。
原文摘要 · Abstract (English)
Different vision neural networks -- trained to classify, contrast, reconstruct, or match images to text -- should have correspondingly different internal representations. We report that they do not. After training, the top sixteen principal directions of variation inside thirteen modern vision encoders converge to the same sixteen-dimensional geometric object. We call this the cross-architecture substrate and study it with PCA, centred kernel alignment (CKA), and Pang 2026 calibration. The substrate transports across four visual domains (natural photographs, medical CT, satellite, microscopy) at median Procrustes-CKA 0.679, and across eight domains (adding sketches, depth, thermal infrared, astronomy) at 0.604, every pair >0.40. It survives Pang calibration globally (7.4x disc-vs-MAE separation, n=13,394) and locally (4.82-5.30, p<10^{-44}). It is not pixel statistics (0.263), not Gabor features (0.31), not a random projection (0.041), and emerges in the first 10% of training while accuracy keeps climbing. We deliver four applications: a label-free transferability filter beating LogME (3x faster, +0.15 Kendall-tau); a four-way domain detector (99.6% accuracy); a frozen low-shot probe (16 dims beat 768-dim DINOv2 by 3.78pp at N=50 labels per class); and a teacher-free distillation auxiliary matching trained-teacher KD on 33 pairs (7.56pp peak gain at 10% label fraction). The substrate does not cross modalities, does not help cross-paradigm distillation, and does not predict transfer quality (rho=0.08 against transfer accuracy).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。