arXiv:2608.03138cs.CLcs.AI2026-08

用单次生成替代多阶段写作,让AI更高效写出结构严谨的论文引言。

Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning

论文配图:Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning
图 1 · 摘自论文原文
  • 将写作流程内化为单一策略,通过显式阶段标记控制
  • 提升内容连贯性与结构合理性,推理效率更高
  • 适合需要高质量引言生成的研究者和论文撰写工具

利用大语言模型(LLMs)生成严谨的论文引言仍具挑战,需协调背景、研究空白、方法与贡献之间的逻辑关系。现有方法将该过程外化为多阶段提示或代理工作流,成本高且易产生跨阶段偏差。本文提出 StructPO,一种结构感知的策略学习框架,将整个多阶段写作流程内化为单次通过的策略,由显式阶段标记控制。StructPO引入结构感知信用分配,解耦局部阶段质量与全局连贯性;采用精炼引导优化,将修改行为内化至首次生成策略中。实验表明,相比基于工作流的基线,StructPO在语义对齐、结构合理性与推理效率上均有提升,具备跨领域泛化能力,在人类评估中与 GPT-5.1 相当,且在缩放至 Qwen3-32B 时表现优异。结果表明,通过细粒度策略优化内化学术写作流程,是替代昂贵外部编排的有效方案。

原文摘要 · Abstract (English)

Generating a rigorous paper introduction with large language models (LLMs) remains challenging, since it requires coordinating background, gap identification, method and contribution within a coherent narrative. Existing solutions externalize this process as multi-stage prompts or agent workflows which are expensive and vulnerable to cross-stage drift. We propose StructPO, a struct-aware policy learning framework that internalizes the entire multi-stage writing workflow into a single-pass policy controlled by explicit stage tokens. StructPO introduces struct-aware credit assignment to decouple local stage quality from global coherence and refinement-guided optimization to internalize revision behavior into the first-pass policy. Experiments show that StructPO improves semantic alignment, structural rationality and inference efficiency over workflow-based baselines, generalizes to out-of-domain settings, and remains competitive with GPT-5.1 in human evaluation when scaled to Qwen3-32B. These results show that internalizing academic writing workflows through fine-grained policy optimization offers a viable alternative to costly external orchestration.

论文生成策略学习引言写作LLM优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。