压缩生成模型同时保持可解释性,适合边缘设备部署。
Disentangled and Distilled Encoder for Out-of-Distribution Reasoning with Rademacher Guarantees
- 用约束优化实现教师-学生蒸馏,保留潜在变量的解耦特性。
- 理论证明蒸馏过程在Rademacher复杂度下仍保持解耦性。
- 模型体积显著减小,适用于资源受限设备推理。
近期,变分自编码器(VAE)的解耦潜在空间被用于对来自与训练数据不同分布的多标签异常样本进行推理。解耦潜在空间指潜在维度与图像生成因素之间存在一一对应关系。本文提出解耦蒸馏编码器(DDE)框架,在降低异常检测模型尺寸以适应资源受限设备的同时,保持解耦特性。DDE将学生-教师蒸馏建模为带解耦约束的优化问题,并基于Rademacher复杂度建立了蒸馏过程中的解耦性理论保证。实验在NVIDIA平台上验证了压缩模型的有效性。
原文摘要 · Abstract (English)
Recently, the disentangled latent space of a variational autoencoder (VAE) has been used to reason about multi-label out-of-distribution (OOD) test samples that are derived from different distributions than training samples. Disentangled latent space means having one-to-many maps between latent dimensions and generative factors or important characteristics of an image. This paper proposes a disentangled distilled encoder (DDE) framework to decrease the OOD reasoner size for deployment on resource-constrained devices while preserving disentanglement. DDE formalizes student-teacher distillation for model compression as a constrained optimization problem while preserving disentanglement with disentanglement constraints. Theoretical guarantees for disentanglement during distillation based on Rademacher complexity are established. The approach is evaluated empirically by deploying the compressed model on an NVIDIA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。