提出一种新生成模型,用离散隐变量实现半监督学习,性能媲美连续变量模型。
Joint-stochastic-approximation Autoencoders with Application to Semi-supervised Learning
- 通过联合随机近似优化数据对数似然与后验分布偏差
- 在MNIST和SVHN上离散隐变量模型性能接近最优连续模型
- 首次证明离散隐变量可成功用于复杂半监督任务
我们分析现有深度生成模型(如VAE和GAN)发现两个问题:一是处理离散观测和隐变量能力不足;二是其优化目标与数据似然间接相关。为此,我们提出联合随机近似(JSA)自编码器——一类构建深度有向生成模型的新算法,适用于半监督学习。该算法直接最大化数据对数似然,并同时最小化后验分布与推断模型之间的包含型KL散度。我们提供理论分析并开展系列实验,验证其优势:对编码器-解码器结构不匹配具有鲁棒性,能一致处理离散与连续变量。特别地,我们在广泛使用的MNIST和SVHN数据集上实证表明,采用离散隐空间的JSA自编码器在半监督任务中表现与采用连续隐空间的当前最优生成模型相当。据我们所知,这是首个成功将离散隐变量模型应用于挑战性半监督任务的演示。
原文摘要 · Abstract (English)
Our examination of existing deep generative models (DGMs), including VAEs and GANs, reveals two problems. First, their capability in handling discrete observations and latent codes is unsatisfactory, though there are interesting efforts. Second, both VAEs and GANs optimize some criteria that are indirectly related to the data likelihood. To address these problems, we formally present Joint-stochastic-approximation (JSA) autoencoders - a new family of algorithms for building deep directed generative models, with application to semi-supervised learning. The JSA learning algorithm directly maximizes the data log-likelihood and simultaneously minimizes the inclusive KL divergence the between the posteriori and the inference model. We provide theoretical results and conduct a series of experiments to show its superiority such as being robust to structure mismatch between encoder and decoder, consistent handling of both discrete and continuous variables. Particularly we empirically show that JSA autoencoders with discrete latent space achieve comparable performance to other state-of-the-art DGMs with continuous latent space in semi-supervised tasks over the widely adopted datasets - MNIST and SVHN. To the best of our knowledge, this is the first demonstration that discrete latent variable models are successfully applied in the challenging semi-supervised tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。