证明了最大似然估计在多种情形下的学习曲线单调性,解决长期开放问题。
On Learning-Curve Monotonicity for Maximum Likelihood Estimators
- 通过变体GPT-5.2 Pro推导出最大似然估计的单调性证明
- 在高斯向量和伽马分布中首次建立前向KL散度完全单调性
- 适用于参数模型中的典型分布,为理论可靠性提供新依据
学习曲线单调性指算法在任意给定数据分布族下,随着数据量增加平均性能持续提升。本文首次为最大似然估计在一系列良好设定的参数模型中建立了非平凡的单调性保证。对于对数损失下的序列预测,我们证明了未知协方差的高斯向量(无论均值已知或未知)以及未知尺度参数的伽马变量,其前向KL散度具有单调性(事实上是完全单调性)。该高斯情形在前述研究中被明确列为开放问题,甚至在一维情况下仍未解决。此外,我们观察到对于反向KL散度,一个广为人知的技巧可使非常一般的指数族实现单调性。所有结果均通过GPT-5.2 Pro的变体推导得出,人类仅负责提示继续生成并验证与转录其证明。
原文摘要 · Abstract (English)
The property of learning-curve monotonicity, highlighted in a recent series of work by Loog, Mey and Viering, describes algorithms which only improve in average performance given more data, for any underlying data distribution within a given family. We establish the first nontrivial monotonicity guarantees for the maximum likelihood estimator in a variety of well-specified parametric settings. For sequential prediction with log loss, we show monotonicity (in fact complete monotonicity) of the forward KL divergence for Gaussian vectors with unknown covariance and either known or unknown mean, as well as for Gamma variables with unknown scale parameter. The Gaussian setting was explicitly highlighted as open in the aforementioned works, even in dimension 1. Finally we observe that for reverse KL divergence, a folklore trick yields monotonicity for very general exponential families. All results in this paper were derived by variants of GPT-5.2 Pro. Humans did not provide any proof strategies or intermediate arguments, but only prompted the model to continue developing additional results, and verified and transcribed its proofs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。