arXiv:2511.01816cs.LGmath.OC2025-11

用度量学习替代重建目标,实现无需设定秩的张量分解。

No-Rank Tensor Decomposition Using Metric Learning

  • 以三元组损失和正则化驱动,学习语义与物理关系自然体现的距离
  • 在人脸、脑连接、星系等数据上保持语义相关性并实现良好聚类
  • 适合小样本科学场景,提供可解释嵌入,优于传统方法

高维数据的张量分解常难以捕捉语义或物理上有意义的结构,尤其当依赖重构目标和固定秩约束时。本文提出一种基于度量学习的无秩张量分解框架,将重构目标替换为以相似性驱动的优化。通过结合三元组损失与多样性及均匀性正则化,该方法学习到的距离能自然反映语义与物理关系,并具备收敛性与度量性质的理论保证。我们在多种数据集上进行评估,包括人脸识别(LFW、Olivetti)、脑连接(ABIDE)以及模拟物理系统(星系、晶体)。与经典方法(PCA、t-SNE、UMAP)、张量分解(CP、Tucker、t-SVD)及深度学习模型(VAE、DEC、基于Transformer的嵌入)相比,本方法生成的嵌入能有效保留物理与语义相关性,且聚类性能具有竞争力。尽管Transformer在大数据下预测精度更高,本方法在小样本场景仍表现稳健,提供可解释嵌入,适用于训练困难的科学领域。该工作确立度量学习作为张量分析的合理范式,强调物理可解释性与语义相关性,而非像素级重构,为数据稀缺领域提供高效稳健的替代方案。

原文摘要 · Abstract (English)

Tensor decomposition of high-dimensional data often struggles to capture semantically or physically meaningful structures, particularly when relying on reconstruction objectives and fixed-rank constraints. We introduce a no-rank tensor decomposition framework based on metric learning, which replaces reconstruction objectives with a similarity-driven optimization. By combining a triplet loss with diversity and uniformity regularization, the method learns embeddings where distances naturally reflect semantic and physical relationships, supported by theoretical guarantees on convergence and metric properties. We evaluate the approach on diverse datasets, including face recognition (LFW, Olivetti), brain connectivity (ABIDE), and simulated physical systems (galaxies, crystals). In comprehensive comparisons against classical methods (PCA, t-SNE, UMAP), tensor decompositions (CP, Tucker, t-SVD), and deep learning models (VAE, DEC, transformer based embeddings), our method produces embeddings that preserve physically and semantically relevant relationships and achieve competitive clustering performance. While transformers often excel in predictive accuracy on large datasets, our method provides interpretable embeddings and remains effective in small-data regimes where transformer training may be infeasible. This work establishes metric learning as a principled paradigm for tensor analysis, emphasizing physical interpretability and semantic relevance over pixel-level reconstruction, and offering an efficient and robust alternative in data-scarce scientific domains.

张量分解度量学习小样本可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。