提出可识别的离散潜层生成模型,让复杂数据建模更透明可解释。
Deep Discrete Encoders: Identifiable Deep Generative Models for Rich Data with Discrete Latent Layers
- 采用多层二值潜变量的有向图模型,通过层级压缩设计提升可识别性。
- 理论证明潜层规模随深度递减可保证参数一致估计,算法支持指数级潜变量高效计算。
- 适用于主题建模、图像表征与教育测试反应时分析,适合需要可解释性的高风险场景。
在生成式AI时代,具有潜在表示的深度生成模型(DGMs)广受欢迎。尽管其表现优异,但统计特性仍不明确:模型常过度参数化、不可识别且为黑箱,部署于高风险场景时引发担忧。为此,我们提出针对丰富数据类型的可解释深度生成模型——深度离散编码器(DDEs),其为多层二值潜变量的有向图模型。理论上,我们给出DDEs的清晰可识别条件,表明深层潜层规模应逐层递减。可识别性确保参数估计一致性,并启发深度结构的可解释设计。计算上,提出分层非线性谱初始化结合惩罚随机近似EM算法的可扩展估计流程,能高效估计含指数级潜成分的模型。针对高维数据与深架构的大量模拟验证了理论结果,并展示算法卓越性能。将DDEs应用于三种不同真实数据集:进行层次化主题建模、图像表征学习及教育测评中的反应时间建模。
原文摘要 · Abstract (English)
In the era of generative AI, deep generative models (DGMs) with latent representations have gained tremendous popularity. Despite their impressive empirical performance, the statistical properties of these models remain underexplored. DGMs are often overparametrized, non-identifiable, and uninterpretable black boxes, raising serious concerns when deploying them in high-stakes applications. Motivated by this, we propose interpretable deep generative models for rich data types with discrete latent layers, called Deep Discrete Encoders (DDEs). A DDE is a directed graphical model with multiple binary latent layers. Theoretically, we propose transparent identifiability conditions for DDEs, which imply progressively smaller sizes of the latent layers as they go deeper. Identifiability ensures consistent parameter estimation and inspires an interpretable design of the deep architecture. Computationally, we propose a scalable estimation pipeline of a layerwise nonlinear spectral initialization followed by a penalized stochastic approximation EM algorithm. This procedure can efficiently estimate models with exponentially many latent components. Extensive simulation studies for high-dimensional data and deep architectures validate our theoretical results and demonstrate the excellent performance of our algorithms. We apply DDEs to three diverse real datasets with different data types to perform hierarchical topic modeling, image representation learning, and response time modeling in educational testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。