通过统一缩放让所有切片等效有用,经典Sliced-Wasserstein可媲美复杂变体。
Understanding Learning with Sliced-Wasserstein Requires Rethinking Informative Slices
- 对每个一维切片进行统一缩放,使所有切片信息量均衡。
- 在多种任务中,调优后的经典SWD性能优于或持平复杂变体。
- 适合希望简化模型设计的机器学习实践者使用。
Wasserstein距离(WD)因样本与计算复杂度限制了实际应用。Sliced-Wasserstein距离(SWD)通过将分布投影到一维子空间,利用一维分布的闭式解高效计算。但在高维下,多数随机投影因测度集中现象变得无信息。尽管已有若干SWD变体聚焦于‘有信息’切片,却常引入额外复杂性、数值不稳定性,并损害SWD的理论性质。面对现有文献多集中于修改切片分布却面临挑战的现状,本文重新审视经典SWD,提出通过对一维WD进行缩放,使所有切片具有同等信息量。在合理数据假设和‘切片信息量’定义下,该缩放可退化为对整个SWD的单一全局缩放因子。这进一步简化为常见机器学习流程中的标准学习率搜索。我们在多种任务上进行了广泛实验,结果表明:经适当配置的经典SWD常能匹配甚至超越更复杂变体的性能。最终回答:‘对于常见学习任务,Sliced-Wasserstein是否已足够?’
原文摘要 · Abstract (English)
The practical applications of Wasserstein distances (WDs) are constrained by their sample and computational complexities. Sliced-Wasserstein distances (SWDs) provide a workaround by projecting distributions onto one-dimensional subspaces, leveraging the more efficient, closed-form WDs for one-dimensional distributions. However, in high dimensions, most random projections become uninformative due to the concentration of measure phenomenon. Although several SWD variants have been proposed to focus on \textit{informative} slices, they often introduce additional complexity, numerical instability, and compromise desirable theoretical (metric) properties of SWD. Amidst the growing literature that focuses on directly modifying the slicing distribution, which often face challenges, we revisit the classical Sliced-Wasserstein and propose instead to rescale the 1D Wasserstein to make all slices equally informative. Importantly, we show that with an appropriate data assumption and notion of \textit{slice informativeness}, rescaling for all individual slices simplifies to \textbf{a single global scaling factor} on the SWD. This, in turn, translates to the standard learning rate search for gradient-based learning in common machine learning workflows. We perform extensive experiments across various machine learning tasks showing that the classical SWD, when properly configured, can often match or surpass the performance of more complex variants. We then answer the following question: "Is Sliced-Wasserstein all you need for common learning tasks?"
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。