模型规模和数据量增大,能自动学会组合式任务的结构。
Scaling can lead to compositional generalization
- 通过扩大模型和数据规模,实现对组合任务的泛化。
- 在覆盖任务空间的训练下,模型可精准逼近组合结构。
- 隐藏层能线性解码任务成分,适合作为模型组合能力评估指标。
神经网络能否系统性地捕捉离散、组合式任务结构?尽管大规模神经网络表现出强大能力,但其仍存在频繁失败案例,引发对其组合性的质疑。本文研究标准神经网络如何在共享组合结构的任务上泛化。发现仅通过扩大数据量和模型规模,即可实现组合泛化。该现象在不同任务编码下均成立,前提是训练分布充分覆盖任务空间。理论上证明:标准多层感知机只需线性数量的神经元,即可任意精度逼近一类通用的组合任务族。此外,若模型成功组合泛化,则任务构成成分可从隐藏激活中线性解码。该解码能力与文本生成图像模型无法组合已知概念的现象高度相关。
原文摘要 · Abstract (English)
Can neural networks systematically capture discrete, compositional task structure despite their continuous, distributed nature? The impressive capabilities of large-scale neural networks suggest that the answer to this question is yes. However, even for the most capable models, there are still frequent failure cases that raise doubts about their compositionality. Here, we seek to understand what it takes for a standard neural network to generalize over tasks that share compositional structure. We find that simply scaling data and model size leads to compositional generalization. We show that this holds across different task encodings as long as the training distribution sufficiently covers the task space. In line with this finding, we prove that standard multilayer perceptrons can approximate a general class of compositional task families to arbitrary precision using only a linear number of neurons with respect to the number of task modules. Finally, we uncover that if networks successfully compositionally generalize, the constituents of a task can be linearly decoded from their hidden activations. We show that this metric correlates with failures of text-to-image generation models to compose known concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。