arXiv:2603.15335cs.LG2026-03

通过残差置换增强数据,提升模型预测精度。

Data Augmentation via Causal-Residual Bootstrapping

  • 基于独立机制原则,对边际分布模型的残差进行置换。
  • 在线性高斯设定下,增强数据使预测模型准确率提升。
  • 适合需要引入因果知识的数据增强场景。

数据增强通过领域相关的修改现有数据点来融入领域知识。例如,图像可通过不同色调或方向的复制进行增强,从而体现图像可能在这些维度上变化的知识。近期工作(Teshima and Sugiyama)探索了因果知识(如A导致B导致C)在条件独立等价下的整合。本文提出一种适用于加性噪声设定的方法,可融入超出马尔可夫等价类的信息。该方法基于独立机制原则,对基于边缘概率分布模型的残差进行置换。在增强数据上训练的预测模型表现出更高的准确性,本文在线性高斯设定下提供了理论支持。

原文摘要 · Abstract (English)

Data augmentation integrates domain knowledge into a dataset by making domain-informed modifications to existing data points. For example, image data can be augmented by duplicating images in different tints or orientations, thereby incorporating the knowledge that images may vary in these dimensions. Recent work by Teshima and Sugiyama has explored the integration of causal knowledge (e.g, A causes B causes C) up to conditional independence equivalence. We suggest a related approach for settings with additive noise that can incorporate information beyond a Markov equivalence class. The approach, built on the principle of independent mechanisms, permutes the residuals of models built on marginal probability distributions. Predictive models built on our augmented data demonstrate improved accuracy, for which we provide theoretical backing in linear Gaussian settings.

数据增强因果推理残差置换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。