用高质量成果反推推理过程,让模型学会严谨的长文本证据生成。
From Final Artifacts to Trajectories: Retrospective Process Supervision for Evidence-Grounded Long-Form Generation

- 从成品反推潜在推理路径,无需人工标注轨迹数据
- 在长文本生成任务中显著提升事实一致性与证据关联度
- 适合需要高可信度推理的法律、报告等专业场景
轨迹数据对提升大语言模型的智能体能力愈发重要。然而,在开放性任务中,由于缺乏唯一正确答案且标注成本高昂,获取大规模轨迹数据极为困难。本文提出 RetroGen——一种自迭代的回溯式过程监督框架。核心观察是:尽管专家推理轨迹稀少,但大量高质量最终成果(如文献综述、分析报告、法律判决)在预训练数据中丰富存在,可视为证据搜寻过程的压缩表征。RetroGen 从这些成果中重构候选隐式轨迹,结合成果与支持证据进行验证,并利用自身成功重构的数据训练模型,无需依赖更强模型生成的轨迹。实验表明,RetroGen 在事实一致性、忠实合成及长文本证据搜寻任务上均有显著提升。
原文摘要 · Abstract (English)
Trajectory data is getting more vital for training large language models for boosting the agentic abilities. Unlike the verifiable domains such as coding or mathematics, scaling trajectory data for open-ended tasks is much more difficult because these tasks lack singular ground truth and are costly to annotate or verify. In this paper, we propose RetroGen, a self-improving framework of retrospective process supervision. Our key observation is that although expert trajectories are scarce, high-quality final artifacts such as literature reviews, analyst reports and legal judgments, are abundant in pre-training data and can be viewed as compressed traces of the evidence-seeking processes that produced them. RetroGen reconstructs candidate latent trajectories from expert artifacts, verifies them against both the artifact and supporting evidence, and trains models on their own successful reconstruction data, without requiring trajectory data from stronger models. Experiments show that RetroGen improves grounding, faithful synthesis, and long-form evidence-seeking agent tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。