通过数据分布特性提升模型系统泛化能力,发现多样性最关键。
Data Distributional Properties As Inductive Bias for Systematic Generalization
- 引入数据多样性、突发性与潜在干预作为归纳偏置
- 多样性使特定属性准确率提升89%绝对值
- 低互信息促进类比推理,适合需要泛化的场景
深度神经网络在系统泛化(SG)上表现不佳。尽管已有研究尝试通过新架构、损失函数或训练方法提升SG,但很少关注训练数据分布特性的作用。本文探究三种数据分布特性作为多模态语言模型的归纳偏置对SG的影响:第一,数据多样性——增加潜在属性取值范围;第二,突发性——在训练中随机限制某些输入的潜在因子取值数量;第三,潜在干预——在训练中随机改变某一潜在因子。实验表明,三者均显著提升SG性能,其中多样性在最受影响属性上带来89%的准确率绝对提升。通过一系列实验,我们发现训练分布中潜在属性间的归一化互信息(NMI)能强预测分布外泛化效果。机制分析表明,较低NMI导致模型表示更平行(即输入特征由并行神经向量编码),这一性质与类比推理能力相关。
原文摘要 · Abstract (English)
Deep neural networks (DNNs) struggle at systematic generalization (SG). Several studies have evaluated the possibility to promote SG through the proposal of novel architectures, loss functions or training methodologies. Few studies, however, have focused on the role of training data properties in promoting SG. In this work, we investigate the impact of certain data distributional properties, as inductive biases for the SG ability of a multi-modal language model. To this end, we study three different properties. First, data diversity, instantiated as an increase in the possible values a latent property in the training distribution may take. Second, burstiness, where we probabilistically restrict the number of possible values of latent factors on particular inputs during training. Third, latent intervention, where a particular latent factor is altered randomly during training. We find that all three factors significantly enhance SG, with diversity contributing an 89% absolute increase in accuracy in the most affected property. Through a series of experiments, we test various hypotheses to understand why these properties promote SG. Finally, we find that Normalized Mutual Information (NMI) between latent attributes in the training distribution is strongly predictive of out-of-distribution generalization. We find that a mechanism by which lower NMI induces SG is in the geometry of representations. In particular, we find that NMI induces more parallelism in neural representations (i.e., input features coded in parallel neural vectors) of the model, a property related to the capacity of reasoning by analogy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。