用强化学习构建科学论文大纲,提升结构与引用一致性
OUTLINEFORGE: Hierarchical Reinforcement Learning with Explicit States for Scientific Writing
- 将论文大纲生成视为分层规划问题,通过结构化操作逐步构建文档
- 在多个指标上优于现有模型,尤其在长程结构连贯性和引用可靠性上提升显著
- 适合需要高质量、结构严谨的科研写作辅助场景
科学论文生成需兼顾全局结构与事实准确性,但当前大语言模型虽局部流畅,常在整体结构、输入覆盖和引用一致性上表现不佳。本文提出一种基于强化学习的框架,将科学大纲构建建模为分层文档结构上的长期规划问题。通过结构化动作建模大纲演化过程,系统可逐步生成完整论文。为促进有效稳定学习,引入两阶段优化:(i) 从部分计划反向重构大纲以保证全局结构一致;(ii) 前向价值引导强化学习,奖励显式建模科学正确性、语篇连贯性和引用忠实度。此外,我们还构建了一个评估基准,涵盖文档规划、输入利用、参考文献忠实度、大纲组织及内容级事实准确性。实验结果表明,本方法在多个强基线模型上实现持续改进,尤其在长距离结构连贯性和引用可靠性方面表现突出。
原文摘要 · Abstract (English)
Scientific paper generation requires document-level planning and factual grounding, but current large language models, despite their strong local fluency, often fail in global structure, input coverage, and citation consistency. We present a reinforcement learning framework that casts scientific outline construction as a long-horizon planning problem over hierarchical document structures. Our approach models edit evolving outlines through structured actions, enabling the system to incrementally build a complete scientific manuscript. To support effective and stabilize learning,we introduce a two-stage optimization procedure consisting of (i) backward outline reconstruction from partial plans to enforce global structural consistency, and (ii) forward value-guided reinforcement learning with rewards explicitly modeling scientific correctness, discourse coherence, and citation fidelity. In addition, We further introduce a benchmark for scientific paper generation that evaluates document planning, input utilization, reference faithfulness, outline organization, and content-level factual accuracy. Our results show consistent improvements over strong neural and LLM baselines, particularly in long-range structural coherence and citation reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。