arXiv:2509.17669cs.CL2025-09被引 1

提出分步生成约束的新方法,让文本生成更可控更真实。

PG-CE: A Progressive Generation Dataset with Constraint Enhancement for Controllable Text Generation

  • 分三步生成:预测类型、构建约束、引导输出
  • 用9万对数据训练,提升生成质量和主题相关性
  • 适合需要精准控制文本风格的场景应用

随着大语言模型的快速发展,可控文本生成(CTG)已成为提升系统可靠性与用户体验的关键技术。针对传统方法的局限,本文提出PG-CE(渐进式生成与约束增强)方法,将CTG任务分解为三个步骤:类型预测、约束构建与引导生成。该方法利用约束生成模型动态构建包含语气、表达风格和主题焦点等多维度约束,以指导输出。实验表明,PG-CE在多个场景下显著提升生成质量,同时保持文本可控性、主题相关性和响应实用性。研究构建了一个包含90,000个约束-文本对的数据集(日常话题与其他话题比例为8:2),有效反映真实应用场景需求。

原文摘要 · Abstract (English)

With the rapid development of Large Language Models (LLMs), Controllable Text Generation (CTG) has become a critical technology for enhancing system reliability and user experience. Addressing the limitations of traditional methods, this paper proposes the PG-CE (Progressive Generation with Constraint Enhancement) approach, which decomposes CTG tasks into three steps: type prediction, constraint construction, and guided generation. This method employs constraint generation models to dynamically build multi-dimensional constraints including tone, expression style, and thematic focus to guide output. Experiments demonstrate that PG-CE significantly improves generation quality across multiple scenarios while maintaining text controllability, thematic relevance, and response practicality. The research developed a dataset containing 90,000 constraint-text pairs (with an 8:2 ratio between daily and other topics), effectively reflecting real-world application requirements.

文本生成可控生成约束学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。