通过可移植表征实现跨域泛化,让模型在未知分布上仍能准确预测。
Abduction-Deduction Entanglement: Domain Generalization via Representation Transplants

- 将预测分解为溯因与演绎两部分,利用源域数据约束目标域的合理推断组合。
- 提出表征移植方法,线性变换表示空间以保留演绎成分、调整溯因信息。
- 适合需要跨域鲁棒性的场景,如医疗诊断、自动驾驶等真实世界应用。
在源分布上训练的预测模型难以泛化到不同的目标分布。对未见数据分布的有效推断必须基于生成源与目标数据的某些因果机制的不变性,但这些结构不变性仅从源数据无法识别。在温和的因果假设下,我们证明目标域的最优预测实际上可通过源分布部分识别。其核心观察是:在任何域中,最优预测可分解为一对溯因与演绎映射——溯因映射从可观测变量推断隐变量(可能为混杂因子),演绎映射则结合可观测与推断量预测标签。大量源数据可确定最优预测,从而约束产生该预测的合法溯因-演绎组合,这种不可识别性称为‘溯因-演绎纠缠’。为此,我们参数化这一受限族,提出‘表征移植’:表示空间中的特定线性变换,在保留演绎成分的同时调整溯因内容。因果机制生成标签的不变性意味着源与目标间存在不变的演绎映射,因此可通过参数化移植搜索合理的目标分布。我们设计学习者-对抗者博弈,理想优化下可收敛至最小最大最优目标预测。实验验证理论,表明该方法在跨域泛化基准上具有竞争力。
原文摘要 · Abstract (English)
Prediction models trained under the source distribution do not generalize well to a different target distribution. A valid inference about an unseen data distribution must be anchored by the invariance of certain causal mechanisms that generate the source and target data, however, these structural invariances are non-identifiable from the source data alone. Under mild causal assumptions about the data, we show that the optimal prediction in the target is in fact partially identifiable by the source distribution. The result rests on a simple observation: In any domain, the optimal prediction can be factorized into what we call a pair of abduction and deduction maps, where the abduction map makes inference about some unobserved variables (possibly confounders) from the observed variables and the deduction map predicts the label using both the observed and inferred quantities. Access to large source data pins down the optimal prediction, thus constrains the valid abduction-deduction ensembles that produce it -- a non-identifiability that we call the abduction-deduction entanglement. To leverage this, we parameterize the constrained family using what we call a representation transplant, that is a specific linear transformation in the representation space that manipulates the abduction content of the representation while retaining the deduction component. Invariance of the causal mechanism generating the label implies existence of an invariant deduction map between source and target. Thus, we can search the space of plausible target distributions via a parametric transplant. We use this scheme in a learner-adversary game that, under an idealistic optimization, provably terminates with the learner having the minimax-optimal target prediction. Evaluations verify the theory, showing that the method is competitive in DG benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。