arXiv:2603.05218cs.AIcs.LG2026-03被引 2

用强化学习训练企业级搜索智能体,性能超越闭源模型。

KARL: Knowledge Agents via Reinforcement Learning

  • 通过多任务强化学习,让智能体在复杂搜索中自主推理并调用工具。
  • 在六个不同搜索场景中表现领先,且对未见过的任务有强泛化能力。
  • 适合需要高精度、低成本的企业知识系统开发者使用。

我们提出一种基于强化学习的系统,用于训练企业级搜索智能体,在多样且难以验证的智能体搜索任务中达到业界最优表现。本文贡献四点:首先,构建KARLBench,涵盖六类搜索场景,包括约束驱动的实体搜索、跨文档报告合成、表格数值推理、全面实体检索、技术文档程序推理及企业内部笔记的事实聚合;其次,证明在异构搜索行为上训练的模型比单一基准优化的模型具有更强泛化能力;第三,设计了一个长周期推理与工具调用结合的智能体合成管道,通过迭代自举生成高质量训练数据;第四,提出一种新的后训练范式,基于大规模离策略强化学习,样本高效、鲁棒性强,天然支持多任务与分布外泛化。相比Claude 4.6和GPT 5.2,KARL在成本-质量、延迟-质量权衡上均呈帕累托最优,包括训练时未见的任务。在充分测试计算资源下,超越最强闭源模型。结果表明,定制合成数据配合多任务强化学习,可实现高效且高性能的知识型智能体。

原文摘要 · Abstract (English)

We present a system for training enterprise search agents via reinforcement learning that achieves state-of-the-art performance across a diverse suite of hard-to-verify agentic search tasks. Our work makes four core contributions. First, we introduce KARLBench, a multi-capability evaluation suite spanning six distinct search regimes, including constraint-driven entity search, cross-document report synthesis, tabular numerical reasoning, exhaustive entity retrieval, procedural reasoning over technical documentation, and fact aggregation over internal enterprise notes. Second, we show that models trained across heterogeneous search behaviors generalize substantially better than those optimized for any single benchmark. Third, we develop an agentic synthesis pipeline that employs long-horizon reasoning and tool use to generate diverse, grounded, and high-quality training data, with iterative bootstrapping from increasingly capable models. Fourth, we propose a new post-training paradigm based on iterative large-batch off-policy RL that is sample efficient, robust to train-inference engine discrepancies, and naturally extends to multi-task training with out-of-distribution generalization. Compared to Claude 4.6 and GPT 5.2, KARL is Pareto-optimal on KARLBench across cost-quality and latency-quality trade-offs, including tasks that were out-of-distribution during training. With sufficient test-time compute, it surpasses the strongest closed models. These results show that tailored synthetic data in combination with multi-task reinforcement learning enables cost-efficient and high-performing knowledge agents for grounded reasoning.

强化学习知识代理企业搜索多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。