arXiv:2507.04504cs.CL2025-07被引 17

用扩散模型生成结构化输出,更稳定可靠。

Unveiling the Potential of Diffusion Large Language Model in Controllable Generation

  • 通过反向推理与全局上下文感知,构建自适应模板框架
  • 在JSON等结构化输出上显著提升格式正确率与内容准确率
  • 适合需要高可靠性输出的智能体、函数调用等场景

可控生成是自然语言处理中的基础任务,广泛应用于函数调用与智能体通信。然而,当前最先进的自回归大语言模型在生成结构化输出时仍不可靠。受新型基于扩散的大语言模型(dLLM)启发,我们发现其架构差异,尤其是全局信息共享机制,可能是突破可控生成的关键。为此,提出自适应模式支架($S^3$)框架,利用dLLM的逆向推理能力与全局上下文感知,直接在输出上下文中初始化结构模板,实现对结构化输出(如JSON)的稳定生成。相比复杂的提示优化,$S^3$方法更具鲁棒性与通用性。实验表明,该方法在结构遵循性、内容忠实度和语义一致性方面显著提升了dLLM的可控生成能力,为语言模型在可控生成任务中的部署提供了新视角与实用路径。

原文摘要 · Abstract (English)

Controllable generation is a fundamental task in NLP with many applications, providing a basis for function calling to agentic communication. However, even state-of-the-art autoregressive Large Language Models (LLMs) today exhibit unreliability when required to generate structured output. Inspired by the current new diffusion-based large language models (dLLM), we realize that the architectural difference, especially the global information-sharing mechanism for language modeling, may be the key to unlock next-level controllable generation. To explore the possibility, we propose Self-adaptive Schema Scaffolding ($S^3$), a novel framework that enables dLLM to stably generate reliable structured outputs (e.g., JSON) by utilizing its innate reverse reasoning capability and global context awareness. $S^3$ initiates a schematic template directly in the output context as a starting state for dLLM, offering a more robust and general method than intricate prompt optimization. Experiments demonstrate that our method substantially unlocks the dLLM's potential in controllable generation in terms of structure adherence, content fidelity, and faithfulness. These results establish new perspectives and practical pathways for deploying language models in controllable generation tasks.

扩散模型结构化生成可控生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。