用拓扑同胚统一不同模态的潜在表示,实现跨域通用建模。
Universal Latent Homeomorphic Manifolds: A Framework for Cross-Domain Representation Unification
- 通过连续双射映射将语义与观测数据统一到同一潜在流形中。
- 在5%像素下恢复人脸图像,跨域分类准确率达86.73%。
- 为大模型拆解提供数学基础,适合跨域学习研究者。
我们提出通用潜在同胚流形(ULHM)框架,将语义表示(如人类描述、诊断标签)与观测驱动的机器表示(如像素强度、传感器读数)统一到单一潜在结构中。尽管来源路径不同,二者均捕捉同一底层现实。我们以拓扑同胚(连续双射且保持结构)作为判定不同语义-观测对潜在流形能否严格统一的数学标准。该准则为三项关键应用提供理论保障:(1) 语义引导的稀疏观测恢复,(2) 跨域迁移学习中的结构兼容性验证,(3) 通过语义到观测空间的有效迁移实现零样本组合学习。框架通过条件变分推断学习连续流形间变换,避免脆弱的点对点映射。我们设计了信任度、连续性及Wasserstein距离等可验证算法,从有限样本中实证验证同胚结构。实验表明:(1) 仅用CelebA 5%像素即可完成稀疏图像恢复,MNIST在多稀疏度下重建成功;(2) 无需重训练,实现从MNIST到Fashion-MNIST的跨域分类,准确率达86.73%;(3) 在未见类别上实现78.76%的零样本分类准确率。关键在于,同胚准则可判断不同语义-观测对是否共享兼容的潜在结构,从而实现有原则的统一为通用表示,并为将通用基础模型分解为领域特定组件提供数学基础。
原文摘要 · Abstract (English)
We present the Universal Latent Homeomorphic Manifold (ULHM), a framework that unifies semantic representations (e.g., human descriptions, diagnostic labels) and observation-driven machine representations (e.g., pixel intensities, sensor readings) into a single latent structure. Despite originating from fundamentally different pathways, both modalities capture the same underlying reality. We establish \emph{homeomorphism}, a continuous bijection preserving topological structure, as the mathematical criterion for determining when latent manifolds induced by different semantic-observation pairs can be rigorously unified. This criterion provides theoretical guarantees for three critical applications: (1) semantic-guided sparse recovery from incomplete observations, (2) cross-domain transfer learning with verified structural compatibility, and (3) zero-shot compositional learning via valid transfer from semantic to observation space. Our framework learns continuous manifold-to-manifold transformations through conditional variational inference, avoiding brittle point-to-point mappings. We develop practical verification algorithms, including trust, continuity, and Wasserstein distance metrics, that empirically validate homeomorphic structure from finite samples. Experiments demonstrate: (1) sparse image recovery from 5% of CelebA pixels and MNIST digit reconstruction at multiple sparsity levels, (2) cross-domain classifier transfer achieving 86.73% accuracy from MNIST to Fashion-MNIST without retraining, and (3) zero-shot classification on unseen classes achieving 78.76% on CIFAR-10. Critically, the homeomorphism criterion determines when different semantic-observation pairs share compatible latent structure, enabling principled unification into universal representations and providing a mathematical foundation for decomposing general foundation models into domain-specific components.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。