提升生成模型效率,让图像生成更快更省资源。
Fast & Efficient Normalizing Flows and Applications of Image Generative Models
- 提出六项创新优化归一化流架构,显著提升计算效率
- 实现快速并行反向推导与高效反向传播算法,加速训练
- 应用于农业质检、地质测绘、自动驾驶隐私保护等真实场景
本论文在两大方向作出新贡献:一是提升生成模型(尤其是归一化流)的效率;二是将生成模型应用于解决实际计算机视觉问题。第一部分提出六项关键改进:1)设计可逆3×3卷积层,给出可逆性的数学充要条件;2)引入更高效的四重耦合层;3)提出k×k卷积层的快速并行反演算法;4)设计卷积逆运算的快速反向传播算法;5)在逆流(Inverse-Flow)中使用卷积逆进行前向传播,并结合所提反向传播算法训练;6)提出轻量级超分辨率模型Affine-StableSR,利用预训练权重和归一化流层,在减少参数量的同时保持性能。第二部分:1)基于条件GAN构建农产品质量自动评估系统,有效应对类别不平衡、数据稀缺与标注难题,实现种子纯度检测高精度;2)提出无监督地质制图框架,采用堆叠自编码器降维,特征提取优于传统方法;3)提出一种基于人脸检测与图像修复的自动驾驶数据集隐私保护方法;4)利用基于Stable Diffusion的图像修复替换检测到的人脸与车牌,推进隐私保护与伦理实践;5)适配扩散模型用于艺术修复,通过统一微调有效处理多种退化类型。
原文摘要 · Abstract (English)
This thesis presents novel contributions in two primary areas: advancing the efficiency of generative models, particularly normalizing flows, and applying generative models to solve real-world computer vision challenges. The first part introduce significant improvements to normalizing flow architectures through six key innovations: 1) Development of invertible 3x3 Convolution layers with mathematically proven necessary and sufficient conditions for invertibility, (2) introduction of a more efficient Quad-coupling layer, 3) Design of a fast and efficient parallel inversion algorithm for kxk convolutional layers, 4) Fast & efficient backpropagation algorithm for inverse of convolution, 5) Using inverse of convolution, in Inverse-Flow, for the forward pass and training it using proposed backpropagation algorithm, and 6) Affine-StableSR, a compact and efficient super-resolution model that leverages pre-trained weights and Normalizing Flow layers to reduce parameter count while maintaining performance. The second part: 1) An automated quality assessment system for agricultural produce using Conditional GANs to address class imbalance, data scarcity and annotation challenges, achieving good accuracy in seed purity testing; 2) An unsupervised geological mapping framework utilizing stacked autoencoders for dimensionality reduction, showing improved feature extraction compared to conventional methods; 3) We proposed a privacy preserving method for autonomous driving datasets using on face detection and image inpainting; 4) Utilizing Stable Diffusion based image inpainting for replacing the detected face and license plate to advancing privacy-preserving techniques and ethical considerations in the field.; and 5) An adapted diffusion model for art restoration that effectively handles multiple types of degradation through unified fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。