arXiv:2506.01414cs.LGcs.IT2025-06TPAMI被引 1

通过星云锚点优化潜在空间,提升多模态任务性能

Self-supervised Latent Space Optimization with Nebula Variational Coding

  • 引入星云锚点引导潜在变量聚类,实现结构化嵌入
  • 自监督度量学习使聚类边界更清晰,提升分类等任务效果
  • 适用于图像、点云、文本等多种数据,通用性强

深度学习模型通过逐层处理数据并生成中间特征(潜在特征)。本文提出一种变分推断模型,旨在通过概率建模优化潜在流形,以提升分类、分割、补全和重建等任务的性能。方法在潜在空间中引入额外变量——星云锚点(nebula anchors),引导潜在变量形成聚类结构。为防止锚点间相互聚集,采用变分约束,使每个锚点周围的潜在特征服从高斯分布,构建出称为星云变分编码(Nebula Variational Coding, NVC)的生成模型。由于每个潜在特征可被标注为距离最近的锚点,进一步提出自监督度量学习,增强不同聚类间的分离性。实验表明,该方法可适配多种架构,应用于文本序列、图像、3D点云及体素数据,有效适应训练数据的语义结构,如样本类别标签,显著提升各类任务表现。

原文摘要 · Abstract (English)

Deep learning approaches process data in a layer-by-layer way with intermediate (or latent) features. We aim at designing a general solution to optimize the latent manifolds to improve the performance on classification, segmentation, completion and/or reconstruction through probabilistic models. This paper proposes a variational inference model which leads to a clustered embedding. We introduce additional variables in the latent space, called \textbf{nebula anchors}, that guide the latent variables to form clusters during training. To prevent the anchors from clustering among themselves, we employ the variational constraint that enforces the latent features within an anchor to form a Gaussian distribution, resulting in a generative model we refer as Nebula Variational Coding (NVC). Since each latent feature can be labeled with the closest anchor, we also propose to apply metric learning in a self-supervised way to make the separation between clusters more explicit. As a consequence, the latent variables of our variational coder form clusters which adapt to the generated semantic of the training data, \textit{e.g.} the categorical labels of each sample. We demonstrate experimentally that it can be used within different architectures designed to solve different problems including text sequence, images, 3D point clouds and volumetric data, validating the advantage of our proposed method.

潜在空间优化自监督学习变分编码聚类结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。