用广义信息散度改进估计风险上限,适用更广分布。
Bounds on the Excess Minimum Risk via Generalized Information Divergence Measures
- 用Rényi和α-Jensen-Shannon散度推广已有风险界
- 在非恒定亚高斯条件下仍有效,且部分情形更紧
- 适合关注统计推断与学习泛化理论的研究者
给定按顺序构成马尔可夫链的有限维随机向量 $Y$, $X$, $Z$(即 $Y \to X \to Z$),本文基于广义信息散度度量推导了估计 $Y$ 时最小风险超出量的上界。其中 $Y$ 是待估计的目标向量,从观测特征向量 $X$ 或其随机退化版本 $Z$ 中估计。超出最小风险定义为从 $X$ 与从 $Z$ 估计 $Y$ 的最小期望损失之差。我们提出一族广义界,推广了 Györfi 等(2023)基于互信息的界,使用 Rényi 散度、$\alpha$-Jensen-Shannon 散度及 Sibson 互信息。这些界类似于 Modak 等(2021)和 Aminian 等(2024)针对学习算法泛化误差的界,但无需假设亚高斯参数恒定,因而适用于更广泛的联合分布。通过数值示例,在常数与非常数亚高斯性假设下均表明,特定 $\alpha$ 参数范围内,基于广义散度的界优于互信息界。
原文摘要 · Abstract (English)
Given finite-dimensional random vectors $Y$, $X$, and $Z$ that form a Markov chain in that order (i.e., $Y \to X \to Z$), we derive upper bounds on the excess minimum risk using generalized information divergence measures. Here, $Y$ is a target vector to be estimated from an observed feature vector $X$ or its stochastically degraded version $Z$. The excess minimum risk is defined as the difference between the minimum expected loss in estimating $Y$ from $X$ and from $Z$. We present a family of bounds that generalize the mutual information based bound of Györfi et al. (2023), using the Rényi and $α$-Jensen-Shannon divergences, as well as Sibson's mutual information. Our bounds are similar to those developed by Modak et al. (2021) and Aminian et al. (2024) for the generalization error of learning algorithms. However, unlike these works, our bounds do not require the sub-Gaussian parameter to be constant and therefore apply to a broader class of joint distributions over $Y$, $X$, and $Z$. We also provide numerical examples under both constant and non-constant sub-Gaussianity assumptions, illustrating that our generalized divergence based bounds can be tighter than the one based on mutual information for certain regimes of the parameter $α$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。