提出可验证的流式批处理方法,提升最优传输计算效率与可靠性。
ForgettingOT: Certified Speculative Batching from Sinkhorn's Projective Forgetting
- 基于Sinkhorn算法的投影遗忘机制,实现流式问题的可验证批处理。
- 实测速度提升1.42倍至3.55倍,且在10^-3精度下无误差违反。
- 适用于需要高可靠性的大规模最优传输任务,如生成模型训练。
正向双边际熵正则最优传输通过非线性、正定、保序且齐次的Sinkhorn映射求解。在对偶缩放规范化后,固定点雅可比矩阵的主特征模λ₂(QP)控制严格修正尾部。投影残差比收敛于该模式,满足容差θ_τ所需的额外认证轮数为log(ρ/θ_τ)/[-logλ₂]+O(1)。ForgettingOT将这一Perron-Frobenius性质转化为相关Sinkhorn问题流的可验证执行器。可计算的投影变差Ω_t及核界约束携带残差;经验证的收缩量q_t给出候选修复深度。窗口定理将这些深度与审计网格转化为对打包工作量、集体轮次、超调和回退的界限。经验尾部估计分配工作但不授权释放;当前实例证书或测量边缘残差才授权释放,普通Sinkhorn作为回退。在15个FP64 A100/OTT-JAX单元上,观测到的商慢模比与λ₂(QP)一致,误差仅9.84×10⁻⁶。在四张A100受控流上,完整执行器比顺序软c-变换预热快1.42倍至3.55倍,30/30配对胜出,无10⁻³边缘容差违规。八张A100支持4096的流实现2.584倍至2.945倍的墙时加速,向量集体轮次减少4.285倍至4.615倍。外层执行器可与保持目标的Sinkhorn加速器组合;若映射变化或目标近似,则需先引入收缩或误差桥以继承修复深度界。
原文摘要 · Abstract (English)
Positive two-marginal entropic optimal transport is solved by a nonlinear, positive, order-preserving, homogeneous Sinkhorn map. After quotienting the dual scaling gauge, we show that the active eigenmode of the fixed-point Jacobian $J_t^\star=QP=P^\star P$, generically $λ_2(QP)$, controls the strict correction tail. The projective-residual ratio converges to this mode, and the additional certified cycles required for tolerance $θ_τ$ scale as $\log(ρ/θ_τ)/[-\logλ_2]+O(1)$. ForgettingOT turns this nonlinear Perron--Frobenius fact into a certified executor for streams of related Sinkhorn problems. A computable projective variation $Ω_t$ in the marginals and kernel bounds the carry residual, while a verified contraction $q_t$ gives candidate repair depth. A window theorem converts these depths and the audit grid into bounds on packed work, collective rounds, overshoot, and fallback. Empirical tail estimates allocate work but never authorize release; current-instance certificates or measured marginal residuals do so, with ordinary Sinkhorn as fallback. On 15 FP64 A100/OTT-JAX cells, the observed quotient slow-mode ratio agrees with $λ_2(QP)$ to $9.84\times10^{-6}$. On controlled four-A100 streams, the complete executor is $1.42\times$--$3.55\times$ faster than sequential soft $c$-transform warm starts, with 30/30 paired wins and no violations of the $10^{-3}$ marginal tolerance. Eight-A100 support-4096 streams give $2.584\times$--$2.945\times$ wall-time speedup and $4.285\times$--$4.615\times$ fewer vector-collective rounds. The outer executor composes with target-preserving Sinkhorn accelerators; a changed map or approximate target needs a contraction or error bridge before inheriting the repair-depth bound.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。