arXiv:2607.00382cs.CV2026-07中稿 · ECCV

针对3D生成模型提出高效压缩方法,显著减小体积且保持形状精度。

Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers

论文配图:Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers
图 1 · 摘自论文原文
  • 基于几何重要性差异,融合结构化剪枝与自适应量化。
  • 在多个主流模型上实现最高66%的模型压缩率。
  • 适合资源受限场景下的3D形状生成应用。

我们提出首个针对图像到形状扩散变换器(DiT)的压缩方法,大幅降低模型规模的同时保持几何保真度。尽管3D形状生成取得显著进展,基于DiT的大模型在资源受限环境下仍计算成本高昂。现有扩散模型压缩策略难以直接迁移至3D生成领域,而以往3D效率研究多聚焦推理速度而非骨干网络压缩。为此,我们构建了专为图像到3D DiT设计的几何感知压缩框架。基于观察:3D DiT层对几何合成的重要性不均,我们引入活力引导框架,集成结构化剪枝、自适应量化与针对性微调。该方法在主流图像到3D模型上实现最高达66%的模型尺寸缩减,同时合成保真度接近全尺寸模型。表明本框架可作为通用插件式方案,广泛适用于多种3D形状生成模型。

原文摘要 · Abstract (English)

We propose the first compression approach for image-to-shape Diffusion Transformers (DiTs) that substantially reduces model size while preserving geometric fidelity. Despite remarkable progress in 3D shape generation, large DiT-based models remain computationally prohibitive in resource-constrained settings. Furthermore, it is difficult to directly transfer existing diffusion model compression strategies developed for different domains to 3D generation, and prior 3D efficiency approaches focus primarily on inference speed rather than backbone compression. To address this limitation, we build a geometry-aware compression framework tailored to image-to-shape DiTs. Guided by the observation that 3D DiT layers exhibit non-uniform importance for geometry synthesis, we introduce a vitality-guided framework integrating structured pruning, adaptive quantization, and targeted fine-tuning. Our method achieves up to 66% model-size reduction across state-of-the-art image-to-3D models while maintaining synthesis fidelity comparable to full-sized counterparts. This highlights the potential of our framework as a plug-and-play solution for efficient 3D shape generation across diverse models.

3D生成扩散模型模型压缩DiT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。