提出新方法量化线性最优传输嵌入的解释能力,评估其在统计分析中的有效性。
Fused Gromov-Wasserstein Variance Decomposition with Linear Optimal Transport
- 基于融合格罗莫夫-沃瑟斯坦距离,分解2-Wasserstein空间中测度的弗雷歇方差
- 实验证明低维线性最优传输嵌入可解释超85%方差,且提升分类准确率
- 适用于需要降维但保留统计结构的图像与高维数据建模场景
Wasserstein距离是概率测度空间中的一类度量,近年来应用广泛。但由于Wasserstein空间的非线性,其上的统计分析复杂。线性最优传输(LOT)提供了一种将测度嵌入欧氏空间的方法,称为LOT嵌入,但会损失部分信息。为评估基于该嵌入的统计推断是否有效,本文提出2-Wasserstein空间中测度集弗雷歇方差的分解方法,可计算嵌入所解释的方差比例。进一步将该分解扩展至融合格罗莫夫-沃瑟斯坦(Fused Gromov-Wasserstein)设置。通过使用MNIST手写数字数据集、IMDB-50000文本数据集和扩散张量MRI图像进行实验,探究了LOT嵌入维度、方差解释比例与机器学习分类器准确率之间的关系。结果表明,低维嵌入可解释超过85%的方差,并在分类任务中保持高准确率。
原文摘要 · Abstract (English)
Wasserstein distances form a family of metrics on spaces of probability measures that have recently seen many applications. However, statistical analysis in these spaces is complex due to the nonlinearity of Wasserstein spaces. One potential solution to this problem is Linear Optimal Transport (LOT). This method allows one to find a Euclidean embedding, called LOT embedding, of measures in some Wasserstein spaces, but some information is lost in this embedding. So, to understand whether statistical analysis relying on LOT embeddings can make valid inferences about original data, it is helpful to quantify how well these embeddings describe that data. To answer this question, we present a decomposition of the Fréchet variance of a set of measures in the 2-Wasserstein space, which allows one to compute the percentage of variance explained by LOT embeddings of those measures. We then extend this decomposition to the Fused Gromov-Wasserstein setting. We also present several experiments that explore the relationship between the dimension of the LOT embedding, the percentage of variance explained by the embedding, and the classification accuracy of machine learning classifiers built on the embedded data. We use the MNIST handwritten digits dataset, IMDB-50000 dataset, and Diffusion Tensor MRI images for these experiments. Our results illustrate the effectiveness of low dimensional LOT embeddings in terms of the percentage of variance explained and the classification accuracy of models built on the embedded data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。