在有限样本下,因果不变性能否提升域适应效果?
How Useful is Causal Invariance for Domain Adaptation in Finite-Sample Settings?

- 基于因果结构识别稳定特征子集,构建候选预测器
- 当候选间目标风险差距足够大时,可实现比仅用目标数据更快的收敛
- 理论揭示了因果知识在小样本域适应中的适用边界
机器学习模型在部署到与训练分布不同的目标分布时性能往往下降。基于因果的域泛化研究发现,跨域共享的因果结构可诱导出不变预测器,即在结构化域偏移下风险稳定的特征子集。然而,这种总体层面的因果不变性在有限样本场景下的实际效用仍不明确。尤其在实践中常仅有少量标注的目标样本,即监督域适应(sDA)场景。本文探讨了完全或部分因果知识在该场景中是否能带来可证明的收益。以线性回归为例,因果知识定义了一组不变或可能不变的特征子集,每个子集可生成一个源数据训练的候选预测器。我们推导出匹配的上下界,表明有限样本增益取决于候选预测器间的目标风险差距,以及源数据估计误差。当这些差距相对于 $n_Q$ 足够大时,自适应聚合方法可逼近最优候选,避免负迁移;反之,若差距过小,则任何算法都无法可靠利用候选集合获得更优的有限样本速率。我们进一步将风险差距与线性结构因果模型中的结构偏移幅度关联,并在真实世界因果基准上验证了理论。
原文摘要 · Abstract (English)
Machine learning models often degrade when they are deployed on a target distribution that differs from the source distributions they were trained on. Recent work in causality-based domain generalization has shown how shared causal structure between domains can induce invariant predictors, e.g., models on a subset of features which have stable risk across structured domain shifts. However, the extent to which such population-level causal invariances can lead to gains in finite-sample settings remains underexplored. In particular, in practice we often have access to a few labeled target samples, a setting called supervised domain adaptation (sDA). In this paper, we explore when (full or partial) causal knowledge can provably improve supervised domain adaptation. As a first step, we study linear regression, where full or partial causal knowledge specifies a collection of invariant or possibly invariant feature subsets, each yielding a source-trained candidate predictor. We derive matching upper and lower bounds showing that finite-sample gains are governed by the target-risk margins separating the candidates, together with the finite-source estimation error. When these margins are sufficiently large relative to $n_Q$, an adaptive aggregation procedure can match the best candidate predictor while avoiding negative transfer relative to target-only learning. On the other hand, when the margins are too small, no algorithm can reliably exploit the candidate collection to obtain faster finite-sample rates. We further connect these margins to structural shift magnitude in linear SCMs and validate the theory on real-world causal benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。