发现Transformer与递归神经网络的新连接,提出可互通的中间模型。
On the Design Space Between Transformers and Recursive Neural Nets
- 通过连续递归网络和神经数据路由,打通RvNN与Transformer的设计鸿沟。
- 新模型在算法任务上表现优于传统RvNN与Transformer,泛化能力更强。
- 适合对模型结构设计、归纳偏置感兴趣的学者深入研究。
本文研究递归神经网络(RvNN)与Transformer两类模型,揭示二者通过近期提出的连续递归神经网络(CRvNN)与神经数据路由(NDR)实现紧密连接。一方面,CRvNN突破传统RvNN的离散结构限制,演变为类似Transformer的架构;另一方面,NDR约束原始Transformer,增强结构归纳偏置,逼近CRvNN形态。两者在算法任务与泛化性能上均显著优于简单形式的RvNN与Transformer。本文探索二者之间的设计空间,形式化其关联,讨论局限性,并提出未来研究方向。
原文摘要 · Abstract (English)
In this paper, we study two classes of models, Recursive Neural Networks (RvNNs) and Transformers, and show that a tight connection between them emerges from the recent development of two recent models - Continuous Recursive Neural Networks (CRvNN) and Neural Data Routers (NDR). On one hand, CRvNN pushes the boundaries of traditional RvNN, relaxing its discrete structure-wise composition and ends up with a Transformer-like structure. On the other hand, NDR constrains the original Transformer to induce better structural inductive bias, ending up with a model that is close to CRvNN. Both models, CRvNN and NDR, show strong performance in algorithmic tasks and generalization in which simpler forms of RvNNs and Transformers fail. We explore these "bridge" models in the design space between RvNNs and Transformers, formalize their tight connections, discuss their limitations, and propose ideas for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。