提出首个通用可迁移性估计方法,解决未知域下模型性能的可信预测问题。
Partial Transportability for Domain Generalization
- 基于因果图和部分识别理论,构建跨域推断的约束参数化框架
- 首次实现对目标域泛化误差的严格边界估计,实验验证其一致性
- 适合关注模型鲁棒性与跨域部署的工业研究者
人工智能中一项基本任务是为在未见域上的预测提供性能保证。实践中,新数据分布存在较大不确定性,导致现有预测器性能波动显著。本文基于部分识别与可迁移性理论,提出新结果:在已知源域数据及数据生成机制假设(以因果图编码)的前提下,可对目标域上某函数值(如分类器的泛化误差)进行边界估计。我们的贡献在于首次提出通用的可迁移性估计技术,通过神经因果模型等参数化方案,编码跨群体推断所需的结构约束。我们证明了该方法的表达力与一致性,并提出基于梯度的优化方案,实现实际可扩展的推断。实验结果验证了方法的有效性。
原文摘要 · Abstract (English)
A fundamental task in AI is providing performance guarantees for predictions made in unseen domains. In practice, there can be substantial uncertainty about the distribution of new data, and corresponding variability in the performance of existing predictors. Building on the theory of partial identification and transportability, this paper introduces new results for bounding the value of a functional of the target distribution, such as the generalization error of a classifier, given data from source domains and assumptions about the data generating mechanisms, encoded in causal diagrams. Our contribution is to provide the first general estimation technique for transportability problems, adapting existing parameterization schemes such Neural Causal Models to encode the structural constraints necessary for cross-population inference. We demonstrate the expressiveness and consistency of this procedure and further propose a gradient-based optimization scheme for making scalable inferences in practice. Our results are corroborated with experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。