针对输入分布偏移下的矩估计问题,提出最优两阶段算法
Minimax Optimal Two-Stage Algorithm For Moment Estimation Under Covariate Shift
- 先在源分布上训练最优估计器,再用似然比重加权校准
- 理论证明达到极小极大下界(对数因子内最优)
- 适合分布偏移场景的稳健估计,尤其适用于真实数据
当训练与测试阶段输入特征分布不一致时,即发生协变量偏移。在该情形下,估计未知函数的矩是一个经典但研究不足的问题,尽管其在现实场景中普遍存在。本文研究了当源分布和目标分布已知时该问题的极小极大下界。为达到极小极大最优(对数因子内),提出一种两阶段算法:首先在源分布上训练函数的最优估计器,然后通过似然比重加权过程校准矩估计。实际中,源与目标分布通常未知,直接估计似然比可能不稳定。为此,我们提出截断版本的估计器,实现双重稳健性,并给出相应的上界。在合成数据上的大量数值实验验证了理论结果,并进一步展示了所提方法的有效性。
原文摘要 · Abstract (English)
Covariate shift occurs when the distribution of input features differs between the training and testing phases. In covariate shift, estimating an unknown function's moment is a classical problem that remains under-explored, despite its common occurrence in real-world scenarios. In this paper, we investigate the minimax lower bound of the problem when the source and target distributions are known. To achieve the minimax optimal bound (up to a logarithmic factor), we propose a two-stage algorithm. Specifically, it first trains an optimal estimator for the function under the source distribution, and then uses a likelihood ratio reweighting procedure to calibrate the moment estimator. In practice, the source and target distributions are typically unknown, and estimating the likelihood ratio may be unstable. To solve this problem, we propose a truncated version of the estimator that ensures double robustness and provide the corresponding upper bound. Extensive numerical studies on synthetic examples confirm our theoretical findings and further illustrate the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。