arXiv:2608.00701stat.MEcs.LG2026-08

提出新方法应对数据分布偏移,提升模型在复杂场景下的泛化能力。

Augmented Inverse Hybrid Weighting: Robust Inference under Deterministic and Random Distribution Shifts

  • 将分布偏移分为系统性偏差与随机扰动,分别用重加权和数据池化处理
  • 在三个真实多中心数据集上,均方误差降低,覆盖度显著提升
  • 适合处理混合偏移场景,尤其适用于医疗等高风险领域的稳健推断

当从一个群体推广证据到另一个群体时,通过重加权使源样本分布匹配目标协变量分布是一种标准做法。该策略适用于可学习的确定性协变量差异,但在存在超出协变量偏移的变化或密度比权重估计不稳定时可能不足。为此,本文提出一种新模型,允许在系统性偏移被纠正后仍存在非系统性变化。残差偏移被建模为概率空间中的随机扰动,无法以可学习方式表示。因此,将系统性偏移视为偏差并由重加权修正,残差随机扰动视为分布不确定性并通过数据集池化处理。在纯随机扰动下,该原则导出增强逆距离加权(AIDW),结合回归增强与方差最优的数据集级池化。对于混合偏移,提出增强逆混合加权(AIHW),在AIDW与标准增强重要性加权之间插值。两种方法通过描述随机扰动强度的‘分布距离’权衡抽样不确定性和分布不确定性。我们建立了方法的渐近性质,并提供插件式调参指导与模型诊断工具。在三个真实世界多站点数据集上的实验表明,相比标准加权基线,均方误差持续下降,且在仅靠协变量偏移调整会低估覆盖度的场景中,经验覆盖度显著改善,证明了所提方法在多样分布偏移场景下的鲁棒性。

原文摘要 · Abstract (English)

Reweighting source samples to match a target covariate distribution is a standard response to distribution shift when generalizing evidence from one population to another. This strategy is well suited to deterministic, learnable covariate discrepancies, but can be insufficient when source--target population differences also contain changes beyond covariate shift or when estimation of the density-ratio weights is unstable. To address this challenge, we introduce a new model that allows non-systematic changes between two population laws after systematic shifts are accounted for. Such residual shift is modeled as random perturbations to the probability space that cannot be represented in a learnable way. In this way, we separate systematic shifts, treated as bias and corrected by reweighting, from residual random perturbations, treated as distributional uncertainty and handled through dataset pooling. Under pure random perturbations, this principle yields Augmented Inverse Distance Weighting (AIDW), which uses regression augmentation and variance-optimal dataset-level pooling. For mixed shifts, we develop Augmented Inverse Hybrid Weighting (AIHW), which interpolates between AIDW and standard augmented importance weighting. Both methods trade off sampling uncertainty and distributional uncertainty via a \emph{distributional distance} that describes the strength of random perturbations. We establish asymptotic properties of the methods, together with plug-in guidance for choosing tuning parameters and model diagnostic tools. Experiments on three real-world multi-site datasets demonstrate consistent reductions in mean-squared error compared with standard weighting baselines, along with substantially improved empirical coverage in settings where covariate-shift adjustment alone undercovers, showing the robustness of the proposed methods across diverse distribution shift scenarios.

分布偏移稳健推断加权方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。