用秩统计量逼近f散度,无需密度比估计,适合高维数据。
Approximating $f$-Divergences with Rank Statistics
- 基于秩直方图构建离散f散度,避免密度比估计
- 估计值随K单调递增,始终为真实散度下界
- 适用于生成建模等场景,支持高维数据
我们提出一种基于秩统计量的f散度近似方法,通过直接分析秩的分布来避免显式密度比估计。对于分辨率参数K,将两个一元分布μ与ν的差异映射为{0,…,K}上的秩直方图,并通过离散f散度度量其与均匀性的偏离,从而得到秩统计量散度估计器。证明了该估计器在K上单调递增,且始终为真实f散度的下界,并在量化域密度比满足弱正则性条件下建立了K→∞时的定量收敛速率。针对高维数据,定义了切片秩统计量f散度,通过对随机投影下的一维构造进行平均,并给出了切片极限的收敛结果。同时推导了有限样本偏差界及估计器的渐近正态性。最后通过基准测试验证了方法性能,并在生成建模实验中展示了其作为学习目标的有效性。
原文摘要 · Abstract (English)
We introduce a rank-statistic approximation of $f$-divergences that avoids explicit density-ratio estimation by working directly with the distribution of ranks. For a resolution parameter $K$, we map the mismatch between two univariate distributions $μ$ and $ν$ to a rank histogram on $\{ 0, \ldots, K\}$ and measure its deviation from uniformity via a discrete $f$-divergence, yielding a rank-statistic divergence estimator. We prove that the resulting estimator of the divergence is monotone in $K$, is always a lower bound of the true $f$-divergence, and we establish quantitative convergence rates for $K\to\infty$ under mild regularity of the quantile-domain density ratio. To handle high-dimensional data, we define the sliced rank-statistic $f$-divergence by averaging the univariate construction over random projections, and we provide convergence results for the sliced limit as well. We also derive finite-sample deviation bounds along with asymptotic normality results for the estimator. Finally, we empirically validate the approach by benchmarking against neural baselines and illustrating its use as a learning objective in generative modeling experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。