arXiv:2602.06540cs.AIcs.CL2026-02被引 1

让小模型也能写深度研究报告,边写边改大纲,效果媲美闭源大模型。

AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research

  • 模仿人类写作过程,边写边动态修改大纲。
  • 在多个评测集上超越主流闭源系统,洞察力提升显著。
  • 适合需要本地部署、保障数据隐私的研究者使用。

生成深度研究报告需要大规模信息获取与洞察驱动的分析,对当前语言模型构成重大挑战。现有方法多采用先规划后撰写的范式,其性能高度依赖初始提纲质量,而构建全面提纲本身又需强推理能力,导致现有系统几乎完全依赖闭源或在线大模型,带来部署障碍及用户数据安全隐私风险。本文提出AgentCPM-Report,一个轻量级但高性能的本地解决方案,包含模拟人类写作流程的框架和一个8B参数的深度研究智能体。该框架采用写作即推理策略(WARP),使模型能在生成过程中动态修订提纲。在此策略下,智能体交替执行基于证据的草稿撰写与推理驱动的深化拓展,协同支持信息获取、知识精炼与迭代提纲演化。为使小型模型具备此能力,我们引入多阶段代理训练策略,包括冷启动、原子技能强化学习与整体流水线强化学习。在DeepResearch Bench、DeepConsult和DeepResearch Gym上的实验表明,AgentCPM-Report在洞察力方面显著优于领先闭源系统。

原文摘要 · Abstract (English)

Generating deep research reports requires large-scale information acquisition and the synthesis of insight-driven analysis, posing a significant challenge for current language models. Most existing approaches follow a plan-then-write paradigm, whose performance heavily depends on the quality of the initial outline. However, constructing a comprehensive outline itself demands strong reasoning ability, causing current deep research systems to rely almost exclusively on closed-source or online large models. This reliance raises practical barriers to deployment and introduces safety and privacy concerns for user-authored data. In this work, we present AgentCPM-Report, a lightweight yet high-performing local solution composed of a framework that mirrors the human writing process and an 8B-parameter deep research agent. Our framework uses a Writing As Reasoning Policy (WARP), which enables models to dynamically revise outlines during report generation. Under this policy, the agent alternates between Evidence-Based Drafting and Reasoning-Driven Deepening, jointly supporting information acquisition, knowledge refinement, and iterative outline evolution. To effectively equip small models with this capability, we introduce a Multi-Stage Agentic Training strategy, consisting of cold-start, atomic skill RL, and holistic pipeline RL. Experiments on DeepResearch Bench, DeepConsult, and DeepResearch Gym demonstrate that AgentCPM-Report outperforms leading closed-source systems, with substantial gains in Insight.

深度研究智能体本地化写作辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。