arXiv:2509.20074cs.CL2025-09

通过挖掘数据中的可组合模板,显著提升模型的泛化能力。

Can Constructions "SCAN" Compositionality ?

  • 从训练数据自动提取可替换槽的伪构造模板
  • 在SCAN数据集上跨分布测试准确率分别达47.8%和20.3%
  • 仅需原数据量40%即可达到良好效果,适合数据稀缺场景

序列到序列模型虽在诸多任务中表现优异,但在组合性与系统泛化方面仍存在困难。我们将其归因于模型未能内化具有约定意义配对的构造,从而无法支持生成性重组。基于此,我们提出一种无监督的伪构造挖掘方法:从训练数据中自动提取含变量槽的模板。在SCAN数据集上,该方法在未见分布的测试中取得显著提升——ADD JUMP准确率达47.8%,AROUND RIGHT达20.3%,且无需架构修改或额外监督。模型在仅使用原训练数据40%的情况下仍具竞争力,体现出强数据效率。结果表明,构建感知预处理是一种替代复杂架构或训练策略的可行路径。

原文摘要 · Abstract (English)

Sequence to Sequence models struggle at compositionality and systematic generalisation even while they excel at many other tasks. We attribute this limitation to their failure to internalise constructions conventionalised form meaning pairings that license productive recombination. Building on these insights, we introduce an unsupervised procedure for mining pseudo-constructions: variable-slot templates automatically extracted from training data. When applied to the SCAN dataset, our method yields large gains out-of-distribution splits: accuracy rises to 47.8 %on ADD JUMP and to 20.3% on AROUND RIGHT without any architectural changes or additional supervision. The model also attains competitive performance with? 40% of the original training data, demonstrating strong data efAciency. Our findings highlight the promise of construction-aware preprocessing as an alternative to heavy architectural or training-regime interventions.

组合性泛化能力数据效率构造挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。