用生成模型可控地制造数据分布偏移,研究模型鲁棒性边界。
Control+Shift: Generating Controllable Distribution Shifts
- 用任意解码器生成模型系统构造不同强度的分布偏移数据。
- 偏移强度上升时模型性能持续下降,人眼几乎察觉不到变化。
- 数据量超阈值后增大数据无益于提升鲁棒性,更强先验更抗偏移。
我们提出一种新方法,利用任意基于解码器的生成模型生成具有分布偏移的真实数据集。该方法系统地构建了不同强度分布偏移的数据集,便于全面分析模型性能退化情况。我们使用这些生成数据集评估多种常用网络的性能,发现即使偏移在人类视觉上几乎不可察觉,模型性能仍随偏移强度增加而持续下降,且数据增强无法缓解此问题。此外,我们发现训练数据集超过一定规模后,继续扩大对模型鲁棒性无影响;更强的归纳偏置可有效提升模型鲁棒性。
原文摘要 · Abstract (English)
We propose a new method for generating realistic datasets with distribution shifts using any decoder-based generative model. Our approach systematically creates datasets with varying intensities of distribution shifts, facilitating a comprehensive analysis of model performance degradation. We then use these generated datasets to evaluate the performance of various commonly used networks and observe a consistent decline in performance with increasing shift intensity, even when the effect is almost perceptually unnoticeable to the human eye. We see this degradation even when using data augmentations. We also find that enlarging the training dataset beyond a certain point has no effect on the robustness and that stronger inductive biases increase robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。