提出HIPIF框架,解决长程任务中上下文干扰问题。
HIPIF: Hierarchical Planning and Information Folding for Long-Horizon LLM Agent Learning

- 通过子目标分解与已完成历史折叠,减少长上下文干扰
- 在三个基准上显著提升长程任务成功率
- 无需额外模型或专家轨迹,适合通用长程智能体
尽管大语言模型在多种任务中表现出色,但在多轮长程任务中性能常下降。现有方法虽通过精细信用分配和分层强化学习缓解稀疏奖励与长期依赖问题,但未直接解决持续增长的历史记忆导致的长上下文干扰,影响全局任务状态追踪与后续推理决策。受人类通过子目标分解和进度总结处理复杂任务的启发,我们提出层次化规划与信息折叠(HIPIF),训练代理以显式子目标组织长程执行,并折叠已完成的子目标历史以降低上下文干扰。为稳定子目标规划与执行,HIPIF结合分层反思与子目标导向过程奖励,指导子目标生成、转换与执行,无需依赖昂贵辅助模型或任务特定专家轨迹。在三个公开可用的智能体基准上的大量实验验证了该方法的有效性。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have demonstrated strong capabilities as autonomous agents across a wide range of tasks, their performance often degrades in multi-turn long-horizon agentic tasks. Existing methods have made progress through fine-grained credit assignment to alleviate long-horizon sparse rewards and hierarchical reinforcement learning to decompose tasks and reduce long-term dependency. However, these methods still do not directly address long-context interference, in which continuously growing histories weaken the agent's ability to track the global task state and impair subsequent reasoning and decision-making. Inspired by the way humans handle complex tasks through subgoal decomposition and completed progress summarization, we propose Hierarchical Planning and Information Folding (HIPIF) for long-horizon LLM agent learning. HIPIF trains the agent end-to-end to organize long-horizon execution around explicit subgoals while folding completed subgoal histories to reduce long-context interference. Furthermore, to stabilize subgoal-based planning and execution, HIPIF combines hierarchical reflection and subgoal-oriented process rewards to guide subgoal generation, transition, and execution, without relying on costly auxiliary models or task-specific expert trajectories. Extensive experiments on three publicly available agentic benchmarks demonstrate the validity of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。