提出平方族概率模型,可高效计算信息量与散度。
Squared families: Searching beyond regular probability models
- 通过线性统计量的平方构造新概率族,处理奇异性后成正则模型。
- 费舍尔信息为双射变换,归一化常数仅需算一次积分。
- 适合需要高效建模与参数估计的统计学习任务。
我们引入平方族,即通过对统计量的线性变换进行平方得到的概率密度族。尽管平方族具有奇异性,但其奇异性易于处理,使该族成为正则模型。处理奇异性后,平方族具备诸多便利性质:其费舍尔信息是源于Bregman生成器的海森度量的共形变换,而该生成器即为归一化常数,进而产生家族上的统计散度。归一化常数具有关键的参数-积分分解特性,只需计算一次与参数无关的积分,即可获得族中所有归一化常数,这不同于指数族。此外,仅需计算一个核心核积分,即可获得费舍尔信息、统计散度与归一化常数。我们进一步揭示平方族在更广泛的g族中的特殊性:去除特殊奇异性后,只有正齐次族与指数族满足费舍尔信息为海森度量的共形变换,且生成器仅通过归一化常数依赖参数。偶数次单项式族也具有参数-积分分解,而指数族不具备。我们研究了平方族在正确指定与错误指定情形下的参数估计与密度估计问题,并利用泛逼近性质证明:当数据量为N、参数量为n时,可对充分光滑的目标密度以$¹O(N^{-1/2})+C n^{-1/4}$速率学习,其中C为某常数。
原文摘要 · Abstract (English)
We introduce squared families, which are families of probability densities obtained by squaring a linear transformation of a statistic. Squared families are singular, however their singularity can easily be handled so that they form regular models. After handling the singularity, squared families possess many convenient properties. Their Fisher information is a conformal transformation of the Hessian metric induced from a Bregman generator. The Bregman generator is the normalising constant, and yields a statistical divergence on the family. The normalising constant admits a helpful parameter-integral factorisation, meaning that only one parameter-independent integral needs to be computed for all normalising constants in the family, unlike in exponential families. Finally, the squared family kernel is the only integral that needs to be computed for the Fisher information, statistical divergence and normalising constant. We then describe how squared families are special in the broader class of $g$-families, which are obtained by applying a sufficiently regular function $g$ to a linear transformation of a statistic. After removing special singularities, positively homogeneous families and exponential families are the only $g$-families for which the Fisher information is a conformal transformation of the Hessian metric, where the generator depends on the parameter only through the normalising constant. Even-order monomial families also admit parameter-integral factorisations, unlike exponential families. We study parameter estimation and density estimation in squared families, in the well-specified and misspecified settings. We use a universal approximation property to show that squared families can learn sufficiently well-behaved target densities at a rate of $\mathcal{O}(N^{-1/2})+C n^{-1/4}$, where $N$ is the number of datapoints, $n$ is the number of parameters, and $C$ is some constant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。