arXiv:2603.20655cs.LGstat.ML2026-03

将LDA推广到非高斯分布,实现更精准的分类校准。

Exponential Family Discriminant Analysis: Generalizing LDA-Style Generative Classification to Non-Gaussian Models

  • 基于指数族构建统一生成框架,自然参数可闭式求解。
  • 分类准确率媲美LDA/QDA,校准误差降低2-6倍。
  • 理论证明其估计量渐近最优,适合高可靠性场景。

我们提出指数族判别分析(EFDA),一种统一的生成框架,将经典线性判别分析(LDA)从高斯假设扩展至任意指数族分布。在各类条件密度属于同一指数族的前提下,EFDA推导出所有自然参数的闭式最大似然估计,并得到关于充分统计量线性的决策规则,恢复了LDA作为特例,同时在原始特征空间中捕捉非线性决策边界。我们证明,在模型正确设定下,EFDA具有渐近校准性和统计效率,并将其推广至K≥2类和多变量数据。通过五种指数族分布(威布尔、伽马、指数、泊松、负二项)的广泛模拟,EFDA的分类精度与LDA、QDA及逻辑回归相当,但期望校准误差(ECE)降低2-6倍;该差距为结构性存在,对所有样本量n和各类别不平衡水平均持续出现,因误设模型始终渐近失校准。进一步证明并实证确认,在正确设定下,EFDA的对数似然比估计量逼近Cramér-Rao界,且是对比中唯一均方误差收敛于零的估计器。完整推导涵盖九种分布。所有四个理论命题均在Lean 4中形式化验证,使用Aristotle(Harmonic)和OpenGauss(Math, Inc.)作为证明生成器,输出由AXLE(Axiom)独立机器检查。

原文摘要 · Abstract (English)

We introduce Exponential Family Discriminant Analysis (EFDA), a unified generative framework that extends classical Linear Discriminant Analysis (LDA) beyond the Gaussian setting to any member of the exponential family. Under the assumption that each class-conditional density belongs to a common exponential family, EFDA derives closed-form maximum-likelihood estimators for all natural parameters and yields a decision rule that is linear in the sufficient statistic, recovering LDA as a special case and capturing nonlinear decision boundaries in the original feature space. We prove that EFDA is asymptotically calibrated and statistically efficient under correct specification, and we generalise it to $K \geq 2$ classes and multivariate data. Through extensive simulation across five exponential-family distributions (Weibull, Gamma, Exponential, Poisson, Negative Binomial), EFDA matches the classification accuracy of LDA, QDA, and logistic regression while reducing Expected Calibration Error (ECE) by $2$-$6\times$, a gap that is structural: it persists for all $n$ and across all class-imbalance levels, because misspecified models remain asymptotically miscalibrated. We further prove and empirically confirm that EFDA's log-odds estimator approaches the Cramér-Rao bound under correct specification, and is the only estimator in our comparison whose mean squared error converges to zero. Complete derivations are provided for nine distributions. Finally, we formally verify all four theoretical propositions in Lean 4, using Aristotle (Harmonic) and OpenGauss (Math, Inc.) as proof generators, with all outputs independently machine-checked by AXLE (Axiom).

分类算法指数族统计建模生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。