arXiv:2508.04492cs.CVcs.AI2025-08

通过稀疏不变的因果差分嵌入,提升图像干预表示的泛化能力。

Learning Robust Intervention Representations with Delta Embeddings

  • 用因果差分嵌入表征干预,仅关注受动作影响的变量变化
  • 在合成与真实场景中,对分布外数据的性能显著超越基线
  • 无需额外监督,可直接从图像对学习鲁棒的因果表示

因果表示学习近年来受到广泛关注,因其有助于提升模型的泛化与鲁棒性。干预图像对的因果表示(文献中称为“可操作反事实”)具有特性:仅改变干预/动作所影响的场景变量。尽管多数研究聚焦于场景变量的因果建模,但对干预本身的表示仍较少关注。本文提出,提升分布外(OOD)鲁棒性的有效策略是关注潜在空间中可操作反事实的表示。具体地,我们引入一种对视觉场景不变且在因果变量上稀疏的因果差分嵌入(Causal Delta Embedding)。基于此,我们提出一种无需额外监督的方法,从图像对中学习因果表示。在因果三元组挑战中的实验表明,因果差分嵌入在分布外设置下表现优异,在合成与真实世界基准上均显著优于基线。

原文摘要 · Abstract (English)

Causal representation learning has attracted significant research interest during the past few years, as a means for improving model generalization and robustness. Causal representations of interventional image pairs (also called ``actionable counterfactuals'' in the literature), have the property that only variables corresponding to scene elements affected by the intervention / action are changed between the start state and the end state. While most work in this area has focused on identifying and representing the variables of the scene under a causal model, fewer efforts have focused on representations of the interventions themselves. In this work, we show that an effective strategy for improving out of distribution (OOD) robustness is to focus on the representation of actionable counterfactuals in the latent space. Specifically, we propose that an intervention can be represented by a Causal Delta Embedding that is invariant to the visual scene and sparse in terms of the causal variables it affects. Leveraging this insight, we propose a method for learning causal representations from image pairs, without any additional supervision. Experiments in the Causal Triplet challenge demonstrate that Causal Delta Embeddings are highly effective in OOD settings, significantly exceeding baseline performance in both synthetic and real-world benchmarks.

因果学习表示学习图像对鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。