arXiv:2502.04907stat.MLcs.LG2025-02被引 1

用共享支撑的离散化方法加速概率测度学习,显著降低计算成本。

Scalable Learning from Probability Measures with Mean Measure Quantization

  • 所有输入测度统一用共享支撑的K点离散化近似
  • 理论证明量化后测度具有一致性,下游任务收敛性有保障
  • 实测性能接近独立量化,但运行时间大幅减少

我们研究数据为概率测度时的统计学习问题。最优传输(OT)常用于比较和操作此类对象,但当测度支撑较大时计算成本过高。本文提出一种基于量化的方法:将所有输入测度近似为共享相同支撑的K点离散测度。理论上建立了量化测度的一致性,并推导了基于量化测度的多个OT下游任务的收敛性保证。在合成与真实数据集上的数值实验表明,该方法性能可媲美独立量化,同时显著降低运行时间。

原文摘要 · Abstract (English)

We consider statistical learning problems in which data are observed as a set of probability measures. Optimal transport (OT) is a popular tool to compare and manipulate such objects, but its computational cost becomes prohibitive when the measures have large support. We study a quantization-based approach in which all input measures are approximated by $K$-point discrete measures sharing a common support. We establish consistency of the resulting quantized measures. We further derive convergence guarantees for several OT-based downstream tasks computed from the quantized measures. Numerical experiments on synthetic and real datasets demonstrate that the proposed approach achieves performance comparable to individual quantization while substantially reducing runtime.

最优传输概率测度量化可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。