通过可控环境研究长时规划能力的形成、塑造与整合机制。
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

- 构建可控多轮环境,系统分析预训练中规划能力如何习得。
- 少量长时序数据即可提升泛化性,错误在长序列中会放大。
- 多教师在线蒸馏可融合通用规划模式,支持跨场景学习。
多轮长时序规划对基础模型智能体至关重要,但其能力如何获得、塑造与整合仍不明确。现有模型基于不可控且不透明的互联网数据训练,难以厘清规划能力的演化过程。为此,我们提出一个统一且受控的多轮环境,可系统研究规划能力在三个阶段的表现:(1)预训练阶段的能力获取。我们考察数据格式、分布与质量,发现通过思维链状态转移建模显式构建世界模型,能显著增强长时序泛化能力;原子技能不足以支持组合泛化,而少量长时序数据即有效;此外,次优轨迹会因误差累积严重损害性能。(2)后训练阶段的能力塑造。通过互信息区分通用规划模式与任务特定知识。对于通用模式,后训练存在三种应用区域:无需、有效、无效;在低质量与长时序条件下,OPD比GRPO具有更广的有效区域,因其提供更一致的更新方向。对于任务知识,从不同知识背景的教师中蒸馏未见流程可能破坏学生原有的世界模型,而无法完全建立新知识。(3)后训练阶段的能力整合。我们证明多教师在线蒸馏(MOPD)可通过收敛至共享规划模式实现能力融合:兼容模式支持跨环境泛化,部分共享模式利于持续学习,而完全冲突模式则引发严重干扰。
原文摘要 · Abstract (English)
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning ability acquisition during pre-training. We study data format, distribution, and quality. Explicit world model construction through CoT state transition modeling yields stronger long-horizon generalization. Atomic skills alone are insufficient for compositional generalization, whereas a litte long-horizon data works. Moreover, suboptimal trajectories severely impair performance because errors amplify over long horizons. (2) Planning ability shaping via GRPO and OPD post-training. Through mutual information, we distinguish general planning patterns from task-specific planning knowledge. For planning patterns, we identify three application regions of post-training: unnecessary, effective, and unsupported. OPD has a broader effective region than GRPO under low-quality and long-horizon settings, as it provides more consistent update directions. For planning knowledge, distilling unseen procedures from a teacher with different knowledge may impair student's prior world modeling without fully establishing new knowledge. (3) Planning ability integration through MOPD post-training. We show that multi-teacher on-policy distillation (MOPD) integrates capabilities by converging to shared planning-pattern across environments. Compatible patterns enable cross-environment generalization, partially shared patterns support continual learning, while completely conflicting patterns cause severe interference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。