arXiv:2604.08124cs.AI2026-04ACL

用分层经验知识提升搜索智能体的推理效率与稳定性

Beyond Stochastic Exploration: What Makes Training Data Valuable for Agentic Search

  • 通过对比分析和多层级聚类提取经验知识
  • 在多个基准上实现显著性能提升,跨任务泛化能力强
  • 适合研究智能体推理与强化学习的学者参考

强化学习(RL)通过整合外部搜索引擎,已成为提升大语言模型(LLMs)推理能力的有效方法。然而,当前基于RL的搜索智能体通常依赖精心设计的结果奖励引导随机探索,导致推理路径低效且训练不稳定。为此,我们提出一种新框架——分层经验(HiExp),通过对比分析和多层级聚类机制,将原始推理轨迹转化为分层经验知识。利用经验对齐训练,有效约束随机探索,使其演变为有策略、以经验驱动的搜索过程。在多个复杂代理搜索与数学推理基准上的大量实验表明,该方法不仅显著提升性能,还展现出强大的跨任务与跨算法泛化能力。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has become an effective approach for advancing the reasoning capabilities of large language models (LLMs) through the strategic integration of external search engines. However, current RL-based search agents often rely on a process of stochastic exploration guided by carefully crafted outcome rewards, leading to inefficient reasoning trajectories and unstable training. To address these issues, we propose a novel framework, Hierarchical Experience (HiExp), to enhance the performance and training stability of search agents. Specifically, we extract empirical knowledge through contrastive analysis and a multi-level clustering mechanism, transforming raw reasoning trajectories into hierarchical experience knowledge. By leveraging experience-aligned training, we effectively regularize stochastic exploration, evolving it into a strategic and experience-driven search process. Extensive evaluations on multiple complex agentic search and mathematical reasoning benchmarks demonstrate that our approach not only achieves substantial performance gains but also exhibits strong cross-task and cross-algorithm generalization.

强化学习智能体推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。