用范畴论框架分析组合泛化,揭示哪些结构识别让未见句子成立。
Compositional Generalization via Structural Identification in a Category-Theoretic Framework
- 将句子建模为从句法地址到词元的函子,通过选择性坍缩诱导柯兰延拓。
- 21种组合泛化类型中,可接受性由不同识别模式决定,残余失败对应无效结构模板。
- 无需训练模型,直接诊断训练语料在特定识别下允许的结构范围。
组合泛化通常通过模型准确率评估。我们转而探讨:在训练中观察到的结构下,哪些结构或词汇识别能使未见的COGS例子成立。句子被表示为从句法地址到词元的函子,选择性坍缩引发柯兰延拓,从而传播已观测到的关联。在21种COGS泛化类型中,可接受性遵循不同的识别模式,而残余失败则分离出未受支持的结构模板。这些数据侧诊断揭示了训练语料在指定识别下所许可的结构,无需训练预测模型。
原文摘要 · Abstract (English)
Compositional generalization is usually evaluated through model accuracy. We instead ask which structural or lexical identifications make held-out COGS examples admissible from the structures observed in training. Sentences are represented as functors from syntactic addresses to lexical tokens, and selective collapses induce Kan extensions that propagate observed associations. Across 21 COGS generalization types, admissibility follows distinct identification profiles, while residual failures separate unsupported structural templates. These data-side diagnoses characterize what the training corpus licenses under specified identifications, without training a predictive model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。