突破传统假设,实现噪声下隐变量的精准识别
Nonparametric Factor Analysis and Beyond
- 提出非参数噪声下的隐变量可识别性理论框架
- 在真实数据中验证了对经济指标的更优估计效果
- 适合处理复杂噪声场景的机器学习与经济学研究者
几乎所有无监督表示学习中的可识别性结果都依赖于加性独立噪声或无噪声假设。本文研究更一般的情形:噪声可为任意形式、依赖隐变量,并在非线性函数中不可逆纠缠。我们提出一个通用框架,用于在非参数噪声设置下识别隐变量。在合适条件下,生成模型即使存在显著噪声,也仅存在子流形层面的不可辨识性;在结构或分布可变性条件下,一般非线性模型的隐变量可识别至平凡不可辨识性。基于该理论框架,我们开发了相应的估计方法,并在多种合成与真实世界数据中验证其有效性。有趣的是,我们对真实国内生产总值增长率的估计,揭示了比官方报告更深入的经济信息。我们期望该框架能为研究人员和实践者处理现实场景中的隐变量提供新视角。
原文摘要 · Abstract (English)
Nearly all identifiability results in unsupervised representation learning inspired by, e.g., independent component analysis, factor analysis, and causal representation learning, rely on assumptions of additive independent noise or noiseless regimes. In contrast, we study the more general case where noise can take arbitrary forms, depend on latent variables, and be non-invertibly entangled within a nonlinear function. We propose a general framework for identifying latent variables in the nonparametric noisy settings. We first show that, under suitable conditions, the generative model is identifiable up to certain submanifold indeterminacies even in the presence of non-negligible noise. Furthermore, under the structural or distributional variability conditions, we prove that latent variables of the general nonlinear models are identifiable up to trivial indeterminacies. Based on the proposed theoretical framework, we have also developed corresponding estimation methods and validated them in various synthetic and real-world settings. Interestingly, our estimate of the true GDP growth from alternative measurements suggests more insightful information on the economies than official reports. We expect our framework to provide new insight into how both researchers and practitioners deal with latent variables in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。