arXiv:2608.30124cs.LGcs.AI2026-08

用张量积注意力提升模型对新组合的泛化能力

TPR-Attention for Combinatorial Generalization

论文配图:TPR-Attention for Combinatorial Generalization
图 1 · 摘自论文原文
  • 引入张量积表示的注意力机制,显式建模组合结构
  • 在组合任务中优于现有架构,显著提升新组合泛化性能
  • 适合需要系统性泛化的场景,如复杂推理与符号学习

系统性泛化仍是深度学习的重大挑战。尤其在组合泛化——对已知变因的新组合进行泛化——方面,人类能轻松应对,而标准神经网络依赖统计相关性,缺乏显式结构表示。本文提出一种新型架构组件:基于张量积表示(TPRs)的注意力机制,将结构归纳偏置嵌入深度学习。通过在组合任务上的受控实验,验证该TPR-注意力机制在组合泛化上优于现有架构。结果表明,将显式组合结构融入神经注意力具有重要价值,为实现系统性泛化的模型指明了可行路径。

原文摘要 · Abstract (English)

Systematic generalization remains a significant challenge in deep learning. In particular, combinatorial generalization - generalizing to new configurations of known factors of variation - is effortless for humans but difficult for standard neural architectures that rely on statistical correlations rather than explicit structural representations. We introduce a new architectural component that embeds structured inductive bias into deep learning: an attention mechanism operating over tensor-product representations (TPRs). Through controlled experiments on compositional tasks, we show that this TPR-attention mechanism outperforms existing architectural components in combinatorial generalization. These results highlight the value of integrating explicit compositional structure into neural attention and point toward a promising path for models capable of systematic generalization.

组合泛化注意力机制结构表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。