arXiv:2507.14811cs.CVcs.AI2025-07

提出可泛化量化框架,让扩散模型在低资源设备上高效运行

SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models

  • 基于结构语义的分段量化,捕捉模型空间差异
  • 双尺度量化保留激活极性不对称性,提升生成图像质量
  • 无需重训练,适配主流部署工具,适合工业落地

扩散模型虽具备卓越生成能力,但计算开销大,难以在资源受限或延迟敏感场景部署。量化能有效降低模型规模与计算成本,其中后训练量化(PTQ)因无需重新训练或训练数据而备受青睐。然而现有扩散模型的PTQ方法多依赖特定架构启发式规则,通用性差且难以融入工业部署流程。为此,我们提出SegQuant——一种统一量化框架,通过自适应融合互补技术提升跨模型泛化能力。该框架包含:(1)基于图结构的分段感知量化策略(SegLinear),捕捉模型结构语义与空间异质性;(2)双尺度量化方案(DualScale),保留极性不对称激活,对维持生成图像视觉保真度至关重要。SegQuant不仅适用于Transformer类扩散模型,更具备广泛适用性,在保持强性能的同时,无缝兼容主流部署工具。

原文摘要 · Abstract (English)

Diffusion models have demonstrated exceptional generative capabilities but are computationally intensive, posing significant challenges for deployment in resource-constrained or latency-sensitive environments. Quantization offers an effective means to reduce model size and computational cost, with post-training quantization (PTQ) being particularly appealing due to its compatibility with pre-trained models without requiring retraining or training data. However, existing PTQ methods for diffusion models often rely on architecture-specific heuristics that limit their generalizability and hinder integration with industrial deployment pipelines. To address these limitations, we propose SegQuant, a unified quantization framework that adaptively combines complementary techniques to enhance cross-model versatility. SegQuant consists of a segment-aware, graph-based quantization strategy (SegLinear) that captures structural semantics and spatial heterogeneity, along with a dual-scale quantization scheme (DualScale) that preserves polarity-asymmetric activations, which is crucial for maintaining visual fidelity in generated outputs. SegQuant is broadly applicable beyond Transformer-based diffusion models, achieving strong performance while ensuring seamless compatibility with mainstream deployment tools.

扩散模型量化高效部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。