arXiv:2501.09163cs.LGcs.AI2025-01NeurIPS被引 8

提出因果视角下的外推理论,仅用一个离域样本也能实现可靠外推。

Towards Understanding Extrapolation: a Causal Lens

  • 基于潜变量模型与因果机制最小变化原则建模外推问题
  • 仅需一个离训练域样本即可实现潜在变量识别与外推
  • 揭示流形光滑性与分布偏移的内在关联,指导实际算法设计

经典分布偏移处理方法通常要求目标分布完全包含在训练分布内。但在实际中,常只有少量目标样本,且可能位于训练支持集之外,这就需要具备外推能力。本文旨在从理论上理解外推何时可行,并提供无需目标分布在支撑集内的原则性方法。我们通过蕴含因果机制最小变化原则的潜变量模型来形式化外推问题,将其转化为潜变量识别问题。在合理假设下,即使仅有一个离域目标样本,也能实现识别,解决最困难情形。理论揭示了底层流形光滑性与偏移特性之间的复杂互动。实验在合成与真实数据上验证了理论发现及其实际意义。

原文摘要 · Abstract (English)

Canonical work handling distribution shifts typically necessitates an entire target distribution that lands inside the training distribution. However, practical scenarios often involve only a handful of target samples, potentially lying outside the training support, which requires the capability of extrapolation. In this work, we aim to provide a theoretical understanding of when extrapolation is possible and offer principled methods to achieve it without requiring an on-support target distribution. To this end, we formulate the extrapolation problem with a latent-variable model that embodies the minimal change principle in causal mechanisms. Under this formulation, we cast the extrapolation problem into a latent-variable identification problem. We provide realistic conditions on shift properties and the estimation objectives that lead to identification even when only one off-support target sample is available, tackling the most challenging scenarios. Our theory reveals the intricate interplay between the underlying manifold's smoothness and the shift properties. We showcase how our theoretical results inform the design of practical adaptation algorithms. Through experiments on both synthetic and real-world data, we validate our theoretical findings and their practical implications.

外推因果推理分布偏移潜变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。