arXiv:2606.30310stat.MLcs.LG2026-06

用累积分布函数提升切片Wasserstein距离的并行计算效率

Highly Data Parallelizable Estimation of the Sliced-Wasserstein Distance Using Cumulative Distribution Functions

论文配图:Highly Data Parallelizable Estimation of the Sliced-Wasserstein Distance Using Cumulative Distribution Functions
图 1 · 摘自论文原文
  • 基于投影数据的累积分布函数设计新估计器,避免排序操作
  • 在高斯混合模型等场景下比传统方法更高效,误差更低
  • 适合联邦学习,可本地计算并聚合分布函数无需传原始数据

切片Wasserstein(SW)距离通过随机投影将高维问题降维为一维最优传输,成为计算高效的替代方案。传统估算方法依赖于基于分位数函数的蒙特卡洛平均,需对投影样本排序且依赖完整数据集。本文提出一类基于投影测度累积分布函数(CDF)的新估计器,避免排序操作,支持大规模数据并行处理。该类估计器包含多个变体,部分由超参数控制方差或平滑性。我们证明其在累积分布函数比分位数函数更易处理的场景(如高斯混合模型)中表现优异,并天然适用于联邦学习——因投影数据的CDF可在本地计算并聚合,无需交换原始样本。

原文摘要 · Abstract (English)

The Sliced Wasserstein (SW) distance has emerged as a computationally attractive alternative to the Wasserstein distance by leveraging one-dimensional optimal transport along random projections. Standard estimators of the SW distance rely on Monte Carlo averages of one-dimensional Wasserstein distances computed via quantile functions, which require sorting projected samples and access to full datasets. In this work, we introduce a new class of estimators for the Sliced Wasserstein distance based on cumulative distribution functions (CDFs) of projected measures, that avoid sorting and scale via massive dataset parallelism. This class includes several estimators, some of them being indexed by hyperparameters controlling their variance or smoothness. We show that they are especially well suited to scenarios in which CDFs are more tractable than quantile functions, such as mixtures of Gaussians, and moreover that they are also naturally compatible with federated learning, since CDFs of projected data can be computed and aggregated locally without requiring the exchange of raw samples.

Wasserstein距离统计估计联邦学习并行计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。