用合成数据提升模型在组合偏移下的鲁棒性,关键在于生成更真实的属性组合样本。
Compositional World Knowledge leads to High Utility Synthetic data
- 通过最小化费希尔散度,强制属性间条件独立,生成更符合真实世界分布的合成数据
- 在CelebA数据集上实现当前最优的最差组准确率(worst-group accuracy)
- 适合关注数据合成、公平性与模型鲁棒性的研究者
机器学习系统在子群体偏移下表现脆弱,尤其当训练中仅观察到部分属性组合时——这种极端情况称为组合偏移。我们探讨:能否通过训练包含所有可能属性组合的合成数据来提升鲁棒性?实验表明,在有限数据上训练的条件扩散模型会学习到错误的底层分布,导致生成的合成样本不真实,无法提升下游性能。为此,我们提出CoInD,通过最小化联合分布与边缘分布之间的费希尔散度,显式建模世界的组合特性,实现条件独立。结果表明,CoInD生成的合成数据高度忠实,使模型在CelebA上的组合偏移任务中达到当前最优的最差组准确率。
原文摘要 · Abstract (English)
Machine learning systems struggle with robustness, under subpopulation shifts. This problem becomes especially pronounced in scenarios where only a subset of attribute combinations is observed during training -a severe form of subpopulation shift, referred as compositional shift. To address this problem, we ask the following question: Can we improve the robustness by training on synthetic data, spanning all possible attribute combinations? We first show that training of conditional diffusion models on limited data lead to incorrect underlying distribution. Therefore, synthetic data sampled from such models will result in unfaithful samples and does not lead to improve performance of downstream machine learning systems. To address this problem, we propose CoInD to reflect the compositional nature of the world by enforcing conditional independence through minimizing Fisher's divergence between joint and marginal distributions. We demonstrate that synthetic data generated by CoInD is faithful and this translates to state-of-the-art worst-group accuracy on compositional shift tasks on CelebA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。