arXiv:2511.11418cs.LGcs.CV2025-11

用最优传输方法量化流匹配模型,2-3比特仍保高质量。

Low-Bit, High-Fidelity: Optimal Transport Quantization for Flow Matching

  • 基于最优传输最小化权重量化后的生成分布差异
  • 2-3比特下视觉质量与潜在空间稳定,优于其他量化方式
  • 适合边缘和嵌入式AI部署,理论与实证兼备

流匹配(Flow Matching, FM)生成模型具备无需模拟训练和确定性采样优势,但实际部署受高精度参数需求制约。本文将基于最优传输(OT)的后训练量化方法应用于FM模型,通过最小化量化前后权重间的2-Wasserstein距离实现压缩,并系统对比了均匀、分段及对数量化方案。理论分析给出了量化导致生成性能下降的上界,实验证明在五个不同复杂度的基准数据集上,该方法可将模型压缩至每参数2-3比特,仍保持良好的图像生成质量与潜在空间稳定性,而其他方法在此条件下失效。这确立了基于最优传输的量化是适用于边缘与嵌入式AI场景的可靠压缩策略。

原文摘要 · Abstract (English)

Flow Matching (FM) generative models offer efficient simulation-free training and deterministic sampling, but their practical deployment is challenged by high-precision parameter requirements. We adapt optimal transport (OT)-based post-training quantization to FM models, minimizing the 2-Wasserstein distance between quantized and original weights, and systematically compare its effectiveness against uniform, piecewise, and logarithmic quantization schemes. Our theoretical analysis provides upper bounds on generative degradation under quantization, and empirical results across five benchmark datasets of varying complexity show that OT-based quantization preserves both visual generation quality and latent space stability down to 2-3 bits per parameter, where alternative methods fail. This establishes OT-based quantization as a principled, effective approach to compress FM generative models for edge and embedded AI applications.

流匹配量化生成模型最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。