提出可分解KL误差的新方法,捕捉变量间高阶交互关系。
A Complete Decomposition of KL Error using Refined Information and Mode Interaction Selection
- 基于信息几何重构对数线性模型,显式分解高阶模式交互
- 算法在合成与真实数据上提升生成与分类性能
- 适合研究高阶依赖建模与能量模型优化的学者
对数线性模型在统计力学和高维统计中长期作为离散变量概率分布学习的基础工具。尽管其广泛应用,现有基于能量模型的方法(如Boltzmann机、马尔可夫图模型)大多仅关注二元关系,忽略变量间的高阶交互结构。本文借助信息几何新工具,重新审视对数线性模型,超越独立分布的一体模式与Boltzmann分布的二体模式,引入高阶模式交互视角,实现对KL误差的完整分解。该分解启发了模式交互项的稀疏选择问题:类似稀疏图选择促进泛化,我们的模型能更高效利用有限数据。我们提出算法MAHGenTa,结合新型蒙特卡洛采样与贪心鲁棒性策略。在合成与真实数据集上,该方法在生成任务中显著提升对数似然,并可轻松拓展至分类等判别任务。
原文摘要 · Abstract (English)
The log-linear model has received a significant amount of theoretical attention in previous decades and remains the fundamental tool used for learning probability distributions over discrete variables. Despite its large popularity in statistical mechanics and high-dimensional statistics, the majority of related energy-based models only focus on the two-variable relationships, such as Boltzmann machines and Markov graphical models. Although these approaches have easier-to-solve structure learning problems and easier-to-optimize parametric distributions, they often ignore the rich structure which exists in the higher-order interactions between different variables. Using more recent tools from the field of information geometry, we revisit the classical formulation of the log-linear model with a focus on higher-order mode interactions, going beyond the 1-body modes of independent distributions and the 2-body modes of Boltzmann distributions. This perspective allows us to define a complete decomposition of the KL error. This then motivates the formulation of a sparse selection problem over the set of possible mode interactions. In the same way as sparse graph selection allows for better generalization, we find that our learned distributions are able to more efficiently use the finite amount of data which is available in practice. We develop an algorithm called MAHGenTa which leverages a novel Monte-Carlo sampling technique for energy-based models alongside a greedy heuristic for incorporating statistical robustness. On both synthetic and real-world datasets, we demonstrate our algorithm's effectiveness in maximizing the log-likelihood for the generative task and also the ease of adaptability to the discriminative task of classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。