arXiv:2603.15158cs.LG2026-03

在代理变量不完美时,仍能精准识别鲁棒预测器。

Point-Identification of a Robust Predictor Under Latent Shift with Imperfect Proxies

  • 用潜在等价类捕捉混淆因子的多重映射关系
  • 仅需跨域混合权重满足秩条件即可点识别
  • 适合存在隐变量偏移的现实数据场景

当分布偏移源于影响协变量和结果的潜在混淆因子时,领域自适应更具挑战性。现有基于代理变量的方法依赖强完备性假设以唯一确定(点识别)鲁棒预测器。完备性要求代理变量包含足够多的潜在混淆因子变化信息。对于不完美的代理变量,从混淆因子到代理分布空间的映射是非单射的,多个潜在混淆因子值可能生成相同的代理分布,从而破坏完备性假设,导致观测数据与多个潜在预测器一致(集合识别)。为此,我们引入潜在等价类(LEC),即诱导相同条件代理分布的一组潜在混淆因子。我们证明,只要多个领域在混合代理诱导的LEC形成鲁棒预测器的方式上存在足够差异,点识别仍可实现。该域多样性条件被形式化为混合权重的跨域秩条件,远弱于完备性假设。我们提出近似贝叶斯主动学习(PQAL)框架,主动查询一组满足秩条件的小而多样化的领域。PQAL可恢复点识别预测器,在合成数据及半合成的dSprites、IHDP、ACS Folktables数据集上表现优于先前方法,且对不同偏移程度具有鲁棒性。

原文摘要 · Abstract (English)

Addressing the domain adaptation problem becomes more challenging when distribution shifts across domains stem from latent confounders that affect both covariates and outcomes. Existing proxy-based approaches that address latent shift rely on a strong completeness assumption to uniquely determine (point-identify) a robust predictor. Completeness requires that proxies have sufficient information about variations in latent confounders. For imperfect proxies the mapping from confounders to the space of proxy distributions is non-injective, and multiple latent confounder values can generate the same proxy distribution. This breaks the completeness assumption and observed data are consistent with multiple potential predictors (set-identified). To address this, we introduce latent equivalent classes (LECs). LECs are defined as groups of latent confounders that induce the same conditional proxy distribution. We show that point-identification for the robust predictor remains achievable as long as multiple domains differ sufficiently in how they mix proxy-induced LECs to form the robust predictor. This domain diversity condition is formalized as a cross-domain rank condition on the mixture weights, which is substantially weaker assumption than completeness. We introduce the Proximal Quasi-Bayesian Active learning (PQAL) framework, which actively queries a small, targeted set of diverse domains that satisfy this rank condition. PQAL can recover the point-identified predictor, demonstrates robustness to varying degrees of shift and outperforms previous methods on synthetic data and semi-synthetic dSprites, IHDP, ACS Folktables datasets.

领域自适应潜在变量鲁棒学习代理变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。