提升神经过程模型在噪声数据下的鲁棒性,避免上下文过拟合。
Robust Neural Processes for Noisy Data
- 用神经过程框架建模函数分布,基于上下文点做预测。
- 注意力机制模型在噪声下易过拟合,性能显著下降。
- 提出简单训练方法,对各类噪声水平均更鲁棒,适合实际噪声场景。
近年来,基于上下文进行自适应预测的模型(即上下文学习)已广泛应用。本文研究此类模型在数据含噪声时的表现。采用神经过程(Neural Processes, NP)框架,该框架能简洁严谨地学习函数分布,预测依赖于一组上下文点。实验发现,在干净数据上表现最佳的模型,与在噪声数据上表现最佳的模型不同:使用注意力机制处理上下文的模型受噪声影响更严重,导致上下文过拟合。为此,本文提出一种简单训练方法,使NP模型在噪声环境下更具鲁棒性。在1D函数和2D图像数据集上的实验表明,该方法在所有噪声水平下均优于其他所有NP模型。
原文摘要 · Abstract (English)
Models that adapt their predictions based on some given contexts, also known as in-context learning, have become ubiquitous in recent years. We propose to study the behavior of such models when data is contaminated by noise. Towards this goal we use the Neural Processes (NP) framework, as a simple and rigorous way to learn a distribution over functions, where predictions are based on a set of context points. Using this framework, we find that the models that perform best on clean data, are different than the models that perform best on noisy data. Specifically, models that process the context using attention, are more severely affected by noise, leading to in-context overfitting. We propose a simple method to train NP models that makes them more robust to noisy data. Experiments on 1D functions and 2D image datasets demonstrate that our method leads to models that outperform all other NP models for all noise levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。