研究不平衡熵正则最优传输的样本复杂度,揭示正则化如何降低数据需求。
Sample complexity of unbalanced entropic OT
- 提出平移不变的对偶形式,分析内在对偶变量的紧性和强凸性
- 得到经验耦合的高概率有限样本界,证明正则化可减少所需样本数
- 适用于机器学习中需高效、稳定运输估计的场景,如生成模型
最优传输(OT)已成为比较概率测度的核心语言,但精确的平衡OT在存在质量缺失、新增或消失的数据时过于僵硬,且在高维下样本复杂度不利。熵正则化和不平衡松弛以互补方式缓解这些问题:熵平滑几何结构、改善统计性质,并支持快速的Sinkhorn类算法;不平衡边缘惩罚则用适配噪声数据的散度项替代严格的守恒约束。本文从最优耦合层面研究熵正则不平衡OT的样本复杂度,提出平移不变的对偶形式,证明内在对偶变量的紧性和强凸性,并将这些几何估计转化为经验耦合的高概率有限样本界。结果阐明了正则化在机器学习应用中的必要性:它弱化了维度诅咒,减少了稳定运输估计所需的样本量,并使估计器兼容可扩展的Sinkhorn类求解器。
原文摘要 · Abstract (English)
Optimal transport (OT) has become a central language for comparing probability measures, but exact balanced OT is often both too rigid for data with missing, created, or destroyed mass and subject to unfavorable high-dimensional sample complexity. Entropic regularization and unbalanced relaxations address these limitations in complementary ways. Entropy smooths the geometry, improves statistical behavior, and enables fast Sinkhorn-type algorithms, while unbalanced marginal penalties replace hard conservation constraints by divergence terms adapted to noisy empirical data. This paper studies the sample complexity of entropic unbalanced OT at the level of the optimal coupling, rather than only the scalar transport value. We develop a translation-invariant dual formulation, prove compactness and strong convexity properties for the intrinsic dual variables, and convert these geometric estimates into high-probability finite-sample bounds for empirical couplings. The results clarify why regularization is a practical necessity in machine learning applications: it softens the curse of dimensionality, reduces the number of samples needed for stable transport estimation, and keeps the resulting estimators compatible with scalable Sinkhorn-type solvers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。