解决源域有隐藏子群体时的无监督域适应问题
Unsupervised Domain Adaptation for Binary Classification with an Unobservable Source Subpopulation
- 通过分布匹配估计不可观测子群体比例
- 理论证明可恢复目标域的背景相关与整体预测模型
- 适合处理存在隐藏分组的现实域适应任务
我们研究一个无监督域适应问题,其中源域由二元标签 $Y$ 和二元背景(或环境)$A$ 定义的子群体组成。重点关注一种挑战性场景:源域中一个子群体不可观测。若忽略该未观测群体,会导致估计偏差和预测性能下降。尽管存在结构化缺失,我们仍能恢复目标域的预测。具体地,我们严格推导出目标域的背景特定及总体预测模型。为实际应用,提出分布匹配方法估计子群体比例。提供估计器渐近行为的理论保证,并建立预测误差上界。在合成与真实数据集上的实验表明,我们的方法优于忽略不可观测子群体的基线方法。
原文摘要 · Abstract (English)
We study an unsupervised domain adaptation problem where the source domain consists of subpopulations defined by the binary label $Y$ and a binary background (or environment) $A$. We focus on a challenging setting in which one such subpopulation in the source domain is unobservable. Naively ignoring this unobserved group can result in biased estimates and degraded predictive performance. Despite this structured missingness, we show that the prediction in the target domain can still be recovered. Specifically, we rigorously derive both background-specific and overall prediction models for the target domain. For practical implementation, we propose the distribution matching method to estimate the subpopulation proportions. We provide theoretical guarantees for the asymptotic behavior of our estimator, and establish an upper bound on the prediction error. Experiments on both synthetic and real-world datasets show that our method outperforms the naive benchmark that does not account for this unobservable source subpopulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。