arXiv:2604.27551cs.LGcs.AI2026-04

通过可控语法生成程序,揭示大模型在分布外场景下的真实泛化能力。

Beyond the Training Distribution: Mapping Generalization Boundaries in Neural Program Synthesis

  • 构建基于特定算术语法的受控环境,系统生成数百万唯一程序。
  • 模型在新语法结构上性能下降超30%,表明其外推能力严重不足。
  • 提升泛化需跨语义与句法空间的多样化训练,适合关注模型鲁棒性的研究者。

大规模Transformer在程序合成基准上表现优异,但其真实泛化能力因数据污染和不透明的训练语料而难以评估。为严格检验模型是否真正泛化而非记忆模板,我们基于领域特定的算术语法构建了严格受控的程序合成环境。通过系统枚举并评估数百万独特程序,我们建立了可解释的句法与语义度量空间,精确映射数据分布,并设计训练与测试划分以隔离特定分布偏移。实验表明,优化密度泛化——即在语义与句法空间中进行多样化采样——可实现稳健的分布外泛化;而支持泛化评估显示,当要求生成句法新颖的程序时,Transformer性能下降超过30%。尽管持续扩大计算资源能改善泛化,但收益遵循严格的对数线性关系。结论指出,稳健泛化需在多流形上最大化训练多样性,且需新型搜索方法突破当前对数线性缩放瓶颈。

原文摘要 · Abstract (English)

Large-scale transformers achieve impressive results on program synthesis benchmarks, yet their true generalization capabilities remain obscured by data contamination and opaque training corpora. To rigorously assess whether models are truly generalizing or merely retrieving memorized templates, we introduce a strictly controlled program synthesis environment based on a domain-specific arithmetic grammar. By systematically enumerating and evaluating millions of unique programs, we construct interpretable syntactic and semantic metric spaces. This allows us to precisely map data distributions and sample train and test splits that isolate specific distributional shifts. Our experiments demonstrate that optimizing density generalization -- through diverse sampling over both semantic and syntactic spaces -- induces robust out-of-distribution generalization. Conversely, evaluating support generalization reveals that transformers severely struggle with extrapolation, experiencing a performance drop of over 30% when forced to generate syntactically novel programs. While steadily scaling up compute improves generalization, the gains follow a strictly log-linear relationship. We conclude that robust generalization requires maximizing training diversity across multiple manifolds, and our findings indicate the necessity for novel search-based approaches to break through current log-linear scaling bottlenecks.

程序合成泛化能力模型评估训练多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。