arXiv:2510.13870cs.CLcs.AI2025-10ACL被引 1

用模板填空法提升扩散语言模型的推理能力

Unlocking the Potential of Diffusion Language Models through Template Infilling

  • 设计模板填空机制,全局约束生成结构
  • 在数学、编程等任务上提升9.40%
  • 适合需要深度推理和多步生成的场景

扩散语言模型(DLMs)作为自回归语言模型的有前途替代方案,其推理策略仍受限于继承自自回归范式的前缀提示。本文提出模板填空(Template Infilling, TI),一种专为DLMs设计的条件生成方法。与传统前缀提示不同,TI在目标响应空间中灵活对齐结构锚点,建立全局蓝图后再填充掩码段。我们在多个基准测试中验证了该方法的有效性,包括数学推理、代码生成和行程规划,相比基线平均提升9.40%。此外,TI在多词元生成场景中展现出额外优势,实现有效加速的同时保持生成质量与鲁棒性。通过施加全局约束,TI最终促进系统2式推理,使模型能在结构化解空间中进行深入思考。

原文摘要 · Abstract (English)

Diffusion Language Models (DLMs) have emerged as a promising alternative to Autoregressive Language Models, yet their inference strategies remain limited to prefix-based prompting inherited from the autoregressive paradigm. In this paper, we propose Template Infilling (TI), a tailored conditioning methodology for DLMs. Unlike conventional prefix prompting, TI flexibly aligns structural anchors across the entire target response space, establishing a global blueprint before filling in the masked segments. We demonstrate the effectiveness of our approach on diverse benchmarks, including mathematical reasoning, code generation, and trip planning, achieving consistent improvements of 9.40% over the baseline. Furthermore, we observe that TI provides additional advantages in multi-token generation settings, enabling effective speedup while maintaining generation quality and robustness. By enforcing these global constraints, TI ultimately facilitates System-2 reasoning, empowering the model to deliberate within a structurally defined solution space.

扩散模型语言模型推理增强模板填空

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。