用高阶归纳类型构建满足代数法则的神经网络,提升组合泛化能力。
Functorial Neural Architectures from Higher Inductive Types
- 将高阶归纳类型编译为神经架构,通过路径构造生成组合结构
- 在环面、楔形圆等空间上,性能比非函子型模型高2-10倍
- 学习到的2-胞射精确修复了克莱因瓶关系下的46%误差
神经网络常能学会任务的组成部分,但在新组合上表现不佳。我们认为这是架构缺陷:解码器仅当遵循任务的代数法则时才能实现组合泛化,即从自由生成序列下降到由这些法则定义的商空间。我们通过将高阶归纳类型(HIT)规范编译为神经架构,使该原则可计算化:基点、路径构造器和2-胞射分别映射为基约束、生成网络、结构拼接和学习的同伦。所得传输解码器天生是严格幺半群函子:拼接词的解码等于独立生成回路段的拼接。相反,我们证明软最大自注意力无法同时满足严格组合性与任意非平凡组合商的下降。在环面、圆楔和克莱因瓶上的实验验证了预测的层级:函子型解码器比非函子型方法高出2–10倍,且学习的2-胞射恰好在涉及克莱因瓶关系的词语上缩小了46%的误差。结果表明,组合泛化应作为架构中的函子结构强制执行,而非仅靠示例学习。
原文摘要 · Abstract (English)
Neural networks often learn the parts of a task but fail on novel combinations of those parts. We argue that this failure is architectural: a decoder generalizes compositionally only when it respects the algebraic laws of the task, i.e. when it descends from freely generated sequences to the quotient determined by those laws. We make this principle constructive by compiling Higher Inductive Type (HIT) specifications into neural architectures. Basepoints, path constructors, and 2-cells are mapped to base constraints, generator networks, structural concatenation, and learned homotopies. The resulting transport decoders are strict monoidal functors by construction: decoding a concatenated word is concatenation of independently generated loop segments. In contrast, we prove that softmax self-attention cannot simultaneously satisfy strict monoidal composition and descent to any non-trivial compositional quotient. Experiments on the torus, wedge of circles, and Klein bottle validate the predicted hierarchy: functorial decoders outperform non-functorial alternatives by $2$--$10\times$, and a learned 2-cell closes a $46\%$ error gap precisely on words exercising the Klein-bottle relation. These results suggest that compositional generalization should be enforced as functorial structure in the architecture, rather than learned from examples alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。