用几何方法揭示视觉与语言模型结构相似但组织方式不同。
On the Spectral Geometry of Cross-Modal Representations: A Functional Map Diagnostic for Multimodal Alignment

- 用图拉普拉斯特征基分析跨模态表示对应关系
- 两模型谱距离仅0.043,结构复杂度相近但基向量几乎不重合
- 提出三个诊断指标评估多模态表示兼容性,适合研究对齐机制
我们采用计算几何中的函数映射框架,研究独立预训练的视觉(DINOv2)与语言(all-MiniLM-L6-v2)编码器之间的跨模态对齐问题。该框架将表示流形间的对应关系建模为图拉普拉斯特征基间的紧凑线性算子。尽管在跨模态检索任务中性能低于Procrustes对齐和相对表示,但该方法揭示了多模态表示的结构性特征:两个编码器的拉普拉斯特征值谱具有高度相似性(归一化谱距离0.043),表明它们捕获的内在复杂度相当。然而,函数映射表现出近乎零的对角主导性(均值低于0.05)和70.15的高正交误差,说明其特征向量基几乎未对齐。我们称此现象为‘谱复杂度-方向间隙’:模型收敛于捕捉结构的能力,却未收敛于组织方式。该间隙定义了谱对齐方法的边界条件,并提出了三个诊断量:对角主导性、正交偏差和拉普拉斯交换误差,用于刻画跨模态表示的兼容性。
原文摘要 · Abstract (English)
We study cross-modal alignment between independently pretrained vision (DINOv2) and language (all-MiniLM-L6-v2) encoders using the functional map framework from computational geometry, which represents correspondence between representation manifolds as a compact linear operator between graph Laplacian eigenbases. While the framework underperforms Procrustes alignment and relative representations for cross-modal retrieval across all supervision budgets, it reveals a structural property of multimodal representations. We find that the Laplacian eigenvalue spectra of the two encoders are quantitatively similar (normalized spectral distance 0.043), indicating that independently trained models develop manifolds of comparable intrinsic complexity. However, the functional map exhibits near-zero diagonal dominance (mean below 0.05) and large orthogonality error (70.15), showing that the eigenvector bases are effectively unaligned. We term this decoupling the spectral complexity--orientation gap: models converge in how much structure they capture but not in how they organize it. This gap defines a boundary condition for spectral alignment methods and motivates three diagnostic quantities : diagonal dominance, orthogonality deviation, and Laplacian commutativity error for characterizing cross-modal representation compatibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。