提出统一框架比较组合模型,揭示神经网络与量子模型在组合泛化上的差异。
Towards a Comparative Framework for Compositional AI Models
- 用范畴论构建无框架依赖的组合模型分析框架
- 两类模型在系统性任务上差距超10%,神经网络更易过拟合
- 通过模块交互解析模型行为,实现可解释性分析
DisCoCirc 框架通过按语法结构组合词单元,构建文本的组合模型。组合性带来两种优势:组合泛化(模型学习整体数据分布的组合规则,从而外推训练分布)与组合可解释性(通过独立分析模块及其组合过程理解模型)。本文以范畴论语言形式化这些概念,并将一系列组合泛化测试适配到该框架。在基于 bAbI 任务扩展的数据集上,对比了基于量子电路与经典神经网络的模型表现。两类模型在生产力与替换性任务上差距小于5%,但在系统性任务上差异至少10%,且在过度泛化任务上趋势不同。总体发现神经模型更易过拟合训练数据。此外,我们展示了如何通过分析模型组件间的交互来解释一个训练好的组合模型的行为。
原文摘要 · Abstract (English)
The DisCoCirc framework for natural language processing allows the construction of compositional models of text, by combining units for individual words together according to the grammatical structure of the text. The compositional nature of a model can give rise to two things: compositional generalisation -- the ability of a model to generalise outside its training distribution by learning compositional rules underpinning the entire data distribution -- and compositional interpretability -- making sense of how the model works by inspecting its modular components in isolation, as well as the processes through which these components are combined. We present these notions in a framework-agnostic way using the language of category theory, and adapt a series of tests for compositional generalisation to this setting. Applying this to the DisCoCirc framework, we consider how well a selection of models can learn to compositionally generalise. We compare both quantum circuit based models, as well as classical neural networks, on a dataset derived from one of the bAbI tasks, extended to test a series of aspects of compositionality. Both architectures score within 5% of one another on the productivity and substitutivity tasks, but differ by at least 10% for the systematicity task, and exhibit different trends on the overgeneralisation tasks. Overall, we find the neural models are more prone to overfitting the Train data. Additionally, we demonstrate how to interpret a compositional model on one of the trained models. By considering how the model components interact with one another, we explain how the model behaves.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。