提出无界密度比估计新方法,解决协变量偏移下的误差控制难题。
Estimating Unbounded Density Ratios: Applications in Error Control under Covariate Shift
- 基于最小二乘与逻辑回归损失,建立无界密度比的误差上界。
- 理论证明源模型可直接泛化至任意目标域,无需损失校正。
- 适用于非参数回归与条件流模型,关键在密度比尾部特性。
密度比是评估两个概率分布相对可能性的重要度量,在统计学和机器学习中有广泛应用。然而,现有密度比估计理论通常依赖严格的正则性条件,主要聚焦于定义域和取值范围有界的密度比函数。本文研究基于最小二乘和逻辑回归损失的密度比估计器,建立了估计误差的上界,达到标准极小极大最优率(对数因子内)。我们的结果适用于定义域和取值范围无界的密度比函数。我们将理论应用于协变量偏移下的非参数回归与条件流模型,识别出密度比的尾部性质是跨域误差控制的关键。我们给出了损失校正不必要的充分条件,并证明源模型可有效泛化到任意合适的目标域。模拟实验支持这些理论发现,表明即使已知真实密度比,源模型仍可能优于损失校正方法得到的模型。
原文摘要 · Abstract (English)
The density ratio is an important metric for evaluating the relative likelihood of two probability distributions, with extensive applications in statistics and machine learning. However, existing estimation theories for density ratios often depend on stringent regularity conditions, mainly focusing on density ratio functions with bounded domains and ranges. In this paper, we study density ratio estimators using loss functions based on least squares and logistic regression. We establish upper bounds on estimation errors with standard minimax optimal rates, up to logarithmic factors. Our results accommodate density ratio functions with unbounded domains and ranges. We apply our results to nonparametric regression and conditional flow models under covariate shift and identify the tail properties of the density ratio as crucial for error control across domains affected by covariate shift. We provide sufficient conditions under which loss correction is unnecessary and demonstrate effective generalization capabilities of a source estimator to any suitable target domain. Our simulation experiments support these theoretical findings, indicating that the source estimator can outperform those derived from loss correction methods, even when the true density ratio is known.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。