用论文元数据自动生成结构严谨、风格一致的学术综述提纲
Meow: End-to-End Outline Writing for Automatic Academic Survey
- 基于论文元数据端到端生成分层提纲,摆脱模板依赖
- 8B模型在结构准确性和风格一致性上表现优异
- 专为自动综述设计,适合研究者快速构建高质量综述框架
随着学术论文数量呈指数增长,利用大语言模型自动开展深度综述已成为必然趋势。提纲撰写作为系统梳理相关工作的关键步骤,在自动综述生成中至关重要。然而,现有方法将提纲写作视为整体流程中的简单步骤,依赖模板化工作流,导致生成的提纲缺乏对主题的深入理解与精细风格表达。为此,我们提出Meow——首个基于元数据的端到端提纲生成框架,能高效生成结构化且忠实于内容的提纲。具体而言,我们将提纲写作建模为从论文元数据生成层级化结构提纲的端到端任务;随后,我们从arXiv、bioRxiv和medRxiv收集高质量综述数据集,并建立系统化的提纲质量评估指标;最后,采用监督微调与强化学习相结合的两阶段训练策略。我们的8B推理模型在结构保真度与风格连贯性方面表现强劲。
原文摘要 · Abstract (English)
As academic paper publication numbers grow exponentially, conducting in-depth surveys with LLMs automatically has become an inevitable trend. Outline writing, which aims to systematically organize related works, is critical for automated survey generation. Yet existing automatic survey methods treat outline writing as mere workflow steps in the overall pipeline. Such template-based workflows produce outlines that lack in-depth understanding of the survey topic and fine-grained styles. To address these limitations, we propose Meow, the first metadata-driven outline writing framework that produces organized and faithful outlines efficiently. Specifically, we first formulate outline writing as an end-to-end task that generates hierarchical structured outlines from paper metadata. We then curate a high-quality dataset of surveys from arXiv, bioRxiv, and medRxiv, and establish systematic evaluation metrics for outline quality assessment. Finally, we employ a two-stage training approach combining supervised fine-tuning and reinforcement learning. Our 8B reasoning model demonstrates strong performance with high structural fidelity and stylistic coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。