让AI通过外部搜索补足知识短板,自动提炼可复用的专业技能。
Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning

- 基于评分标准的强化学习,优化何时搜、如何搜、怎样生成技能
- 在8个专家领域中超越基线方法,提升技能抽象能力
- 适合需要持续进化的AI代理,尤其在专业领域应用
可复用的技能封装了解决真实专业任务所需的程序性知识,为基于大模型的智能体提供了向专家领域自我演进的路径。现有自演化技能方法仅依赖模型参数知识或轨迹数据构建技能,受限于模型已有认知边界。而专业领域的规范与标准流程常超出此边界,难以仅凭模型自身获取。为此,我们提出新框架Search2Skill,能自动识别智能体的能力缺口,搜索外部资源填补,并将检索证据提炼为结构化、可复用的技能。该框架采用基于评分标准的强化学习策略,联合优化搜索时机、搜索方式与技能生成过程。在三个基准的8个专家级领域上实验表明,Search2Skill在流式和离线评估协议下均显著优于搜索增强型与轨迹驱动型基线方法。进一步分析显示,性能提升源于技能抽象而非原始检索内容,且所获技能具备跨模型规模迁移能力。
原文摘要 · Abstract (English)
Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains. Existing self-evolving skill methods construct skills internally from the model's parametric knowledge or trajectories, and are therefore bounded by what the model already knows. However, the domain conventions and standard procedures underlying professional skills often lie beyond this boundary and are hard to elicit from the agent alone. To address this issue, we therefore propose a novel framework, Search2Skill, that automatically identifies the agent's capability gaps, searches external sources to address them, and distills the retrieved evidence into structured, reusable skills. Specifically, Search2Skill is optimized by a rubric-based reinforcement learning scheme that jointly improves when to search, how to search, and how to generate skills. Experiments on eight expert-level domains from three benchmarks show that Search2Skill consistently outperforms both search-augmented and trajectory-based skill-learning baselines under both streaming and held-out evaluation protocols. Further analyses show that the gains arise from skill abstraction rather than raw retrieved evidence, and that the acquired skills transfer across model scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。