arXiv:2605.09485cs.LGstat.ML2026-05被引 1

构建1700个视觉模型的语义嵌入数据集,助力跨模型语义对齐研究。

SEMASIA: A Large-Scale Dataset of Semantically Structured Latent Representations

论文配图:SEMASIA: A Large-Scale Dataset of Semantically Structured Latent Representations
图 1 · 摘自论文原文
  • 收集1700个不同架构与训练配置的视觉模型嵌入,附结构化元数据。
  • 发现模型间嵌入空间存在一致的原型聚类与层次语义邻域结构。
  • 适用于模型可解释性、迁移学习及异构AI系统协同的研究者。

神经网络学习的潜在表示常表现出语义结构,即概念相似性反映在嵌入空间中的几何邻近性。然而,跨模型比较这些空间仍面临挑战:架构、预训练数据、目标函数或随机种子的变化可能导致内容相似但几何不兼容的嵌入。这一潜在空间对齐问题在可解释性、迁移学习、多模态学习、联邦系统和语义通信中至关重要,但进展受限于缺乏大规模、模型多样且元数据丰富的基准。为此,我们提出SEMASIA,一个包含约1700个预训练视觉模型的潜在表示集合,覆盖八个标准图像分类基准。SEMASIA为嵌入提供结构化元数据,涵盖架构、训练策略、预训练来源和模型规模。我们展示了该资源的三项应用:首先,分析单个潜在空间的概念组织,发现模型与数据集间存在一致的原型聚类和层次语义邻域;其次,使用重建误差与下游任务性能基准化监督对齐映射;第三,开展大规模回归分析,探究预训练数据复杂度、专业化程度、迁移学习、增强策略与模型规模如何影响嵌入的几何与探测特性。通过将表征规模与标准化元数据结合,SEMASIA为研究潜在几何、评估对齐方法及开发下一代异构互操作人工智能系统提供了可复现的基础。

原文摘要 · Abstract (English)

Latent representations learned by neural networks often exhibit semantic structure, where concept similarity is reflected by geometric proximity in embedding space. However, comparing such spaces across models remains difficult: changes in architecture, pretraining data, objective, or random seed can yield embeddings with similar content but incompatible geometry. This latent space alignment problem is central to interpretability, transfer and multimodal learning, federated systems, and semantic communication; however, progress remains limited by the lack of large-scale, model-diverse, and metadata-rich benchmarks. To address this gap, we introduce SEMASIA, a large-scale collection of latent representations extracted from approximately 1,700 pretrained vision models across eight standard image-classification benchmarks. SEMASIA pairs embeddings with structured metadata describing architectures, training regimes, pretraining sources, and model scale. We demonstrate three applications of the resource. First, we analyze the conceptual organization of individual latent spaces, showing consistent prototype-like clustering and hierarchical semantic neighborhoods across models and datasets. Second, we benchmark supervised alignment mappings between latent spaces using reconstruction error and downstream task performance. Third, we perform a large-scale regression analysis of how pretraining-data complexity, specialization, transfer learning, augmentation, and model scale relate to geometric and probing properties of embeddings. By coupling representational scale with standardized metadata, SEMASIA provides a reproducible foundation for studying latent geometry, evaluating alignment methods, and developing next-generation heterogeneous and interoperable AI systems.

嵌入对齐视觉模型语义结构基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。