用核空间重构思想学习数据隐含结构,提升降维效果。
Learning Reconstructive Embeddings in Reproducing Kernel Hilbert Spaces via the Representer Theorem
- 在再生核希尔伯特空间中,通过优化表示定理实现样本间线性重构。
- 将高维重构几何映射到低维嵌入空间,保持数据内在结构。
- 适用于复杂数据如癌症分子和物联网入侵检测,效果显著。
针对高维数据表示学习中揭示潜在结构的需求,本文提出基于再生核希尔伯特空间(RKHS)的重建式流形学习新算法。每个样本在RKHS中被表示为其他样本的线性组合,通过优化向量形式的表示定理以实现自表示特性。采用可分离算子值核,使方法适用于向量值数据,同时保持单一标量相似性函数的简洁性。后续的核对齐任务将数据投影至低维隐空间,其格拉姆矩阵旨在匹配高维重构核,从而将RKHS中的自重构几何传递至嵌入空间。因此,该方法扩展了自然数据普遍表现出的自表示性质,结合了核学习理论的经典成果。在模拟数据(同心圆、瑞士卷)和真实数据(癌症分子活性、物联网网络入侵)上的实验验证了方法的实际有效性。
原文摘要 · Abstract (English)
Motivated by the growing interest in representation learning approaches that uncover the latent structure of high-dimensional data, this work proposes new algorithms for reconstruction-based manifold learning within Reproducing-Kernel Hilbert Spaces (RKHS). Each observation is first reconstructed as a linear combination of the other samples in the RKHS, by optimizing a vector form of the Representer Theorem for their autorepresentation property. A separable operator-valued kernel extends the formulation to vector-valued data while retaining the simplicity of a single scalar similarity function. A subsequent kernel-alignment task projects the data into a lower-dimensional latent space whose Gram matrix aims to match the high-dimensional reconstruction kernel, thus transferring the auto-reconstruction geometry of the RKHS to the embedding. Therefore, the proposed algorithms represent an extended approach to the autorepresentation property, exhibited by many natural data, by using and adapting well-known results of Kernel Learning Theory. Numerical experiments on both simulated (concentric circles and swiss-roll) and real (cancer molecular activity and IoT network intrusions) datasets provide empirical evidence of the practical effectiveness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。