arXiv:2603.19954cs.AIcs.CL2026-03中稿 · ICML被引 1

研究Transformer验证长计划的能力,发现其在特定结构下可泛化。

On the Ability of Transformers to Verify Plans

  • 设计新框架C*-RASP,支持序列与词表同步增长下的长度泛化分析
  • 证明在多数经典规划领域中,Transformer能正确验证长计划
  • 揭示影响模型泛化能力的关键结构特性,适合规划与AI理论研究者

Transformer在人工智能规划任务中表现不一,且对泛化何时可能发生缺乏理论理解。本文聚焦解码器仅模型验证给定计划是否正确解决规划实例的能力。针对测试时对象数量(即有效输入词表)增长的通用场景,提出C*-RASP——C-RASP的扩展版本,用于建立Transformer在序列长度与词汇量同时增长下的长度泛化保证。研究识别出一大类经典规划领域,其中Transformer可被证明学习到验证长计划的能力,并揭示显著影响长度泛化解可学习性的结构特性。实证实验验证了理论结果。

原文摘要 · Abstract (English)

Transformers have shown inconsistent success in AI planning tasks, and theoretical understanding of when generalization should be expected has been limited. We take important steps towards addressing this gap by analyzing the ability of decoder-only models to verify whether a given plan correctly solves a given planning instance. To analyse the general setting where the number of objects -- and thus the effective input alphabet -- grows at test time, we introduce C*-RASP, an extension of C-RASP designed to establish length generalization guarantees for transformers under the simultaneous growth in sequence length and vocabulary size. Our results identify a large class of classical planning domains for which transformers can provably learn to verify long plans, and structural properties that significantly affects the learnability of length generalizable solutions. Empirical experiments corroborate our theory.

Transformer规划泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。